Pick questions whose correct answer, source, or safe refusal is known.
Set the operating target.
Describe the live agent and choose the reliability threshold that should trigger investigation.
Define the known answers.
Each row should test one business-critical fact, source, action, or safe refusal. Required phrases use a vertical bar between alternatives.
| # | Canary question | Assertion | Required phrase(s) | Source or tool marker |
|---|
Replay baseline and current outputs.
Paste what the agent returned while healthy and what it returns now. The lab checks required language and source or tool markers.
| # | Question | Healthy baseline | Current output |
|---|
Reliability evidence pack.
Use this result to contain silent drift, brief leadership, or open a vendor support case with reproducible evidence.
No current evidence
Paste current outputs to measure the canary pack.
Canary drift map
Immediate operating response
30-day knowledge reliability protocol
- Day 0: Freeze nonessential changes, name the agent owner, and approve the initial canary pack.
- Days 1–7: Run daily and after every knowledge, instruction, connector, model, permission, or environment change.
- Days 8–14: Add canaries for every real user complaint and every high-cost false answer.
- Days 15–21: Repeat each canary three times to expose inconsistent retrieval, not only total failure.
- Days 22–30: Review pass rate, time to detect, time to restore, false alarms, and unresolved root causes with operations and IT.
