Kiro10d agoContinuous Prompt Evaluation: How We Use LLM Judges and Live Signals to Improve Kiro Agent QualityPrompt behavior is difficult to validate exhaustively. A system prompt operates across combinations of models, tools, codebases, tasks, and users that no test...AIDevTools1 min
Kiro10d agoHow We Learned to Trust an AI Agent to Triage Production IncidentsAt 2:33 AM PDT on a Sunday, an availability alarm fired for a frontier model. Production responses were stalling mid-stream, and monitoring opened a ticket...AIDevTools1 min