Lab · intermediate
Model regression testing
How to run a small, repeatable canary suite after every snapshot rotation so silent regressions surface before production traffic notices.
Difficulty
intermediate
Estimated time
30 min
Output
2 Markdown artifacts
A planning recipe, not a safety validation.
The canary suite catches drift; it does not certify the model. A passing canary is not production readiness, and a failing canary is not a vendor allegation — investigate before escalating.
Goal
Catch a silent regression introduced by a snapshot rotation before downstream users hit it.
When to use this
- Your provider rotates model snapshots without changing the API model name.
- Your automation runs unattended for hours or days at a time.
- You have an evidence brief from a previous selection round and want to keep it honest.
Prerequisites
Test setup
- Freeze a canary suite of 10–25 prompts that exercises the behaviours your application depends on.
- Store reference outputs from the snapshot you originally selected — keep them out of source control if they contain real data.
- Schedule the canary suite to run on a regular cadence (every release, every snapshot bump, or every 24 hours).
- Wire alerting that fires when canary pass rate drops below your floor.
Minimum test set
- 10–25 representative prompts you also ran in the original selection.
- At least one prompt per behaviour your application depends on (recall, format, refusal, modality).
Prompt variants
Vary one dimension at a time so you can attribute behaviour.
- Same prompt, same parameters, two snapshots — reference vs current.
- Same prompt with and without retrieved context to separate model drift from retrieval drift.
Observations to record
Record observations, not scores. The evidence brief stays auditable.
- Pass / fail per prompt against reference output (or rubric, if exact match is not appropriate).
- Latency drift over time.
- Cost drift over time (input/output token shifts can move cost without changing pricing).
- Lifecycle field changes against /reverification.
Failure modes to watch
- Pass rate drops sharply on one prompt category — likely targeted regression.
- Pass rate drops gradually across categories — likely a generic snapshot drift.
- Cost climbs while pass rate stays flat — verbose-output drift.
- The provider deprecates the snapshot mid-run — confirm with the catalogue's lifecycle field.
Stop conditions
Knowing when to stop is part of the test.
- Stop and re-evaluate if pass rate drops below your floor on any release.
- Stop and re-verify if the catalogue's lifecycle field shifts to deprecated for the snapshot under test.
Outputs
- A Markdown evidence delta listing the prompts that regressed.
- An updated decision brief if the regression triggers a re-selection.
Weak test vs stronger test
Weak test
- Run a freeform smoke test ad-hoc and call it a regression check.
- Skip storing reference outputs from the originally-selected snapshot.
- Treat a single failing canary as a vendor issue without investigating.
Stronger test
- Freeze a canary suite of 10–25 prompts covering the behaviours the application depends on.
- Store reference outputs (or rubric criteria) for comparison.
- Schedule the suite on a regular cadence and wire alerting against a documented pass-rate floor.
- Investigate a failing canary against /reverification before escalating.
Observation rubric
Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.
| Dimension | What to look for | What to record |
|---|---|---|
| Per-prompt drift | Whether each canary prompt still meets the reference expectation. | Pass / fail per prompt with rationale, not a rolled-up percentage. |
| Latency drift | Whether latency for the canary set has shifted from the baseline. | Median + tail latency per canary run, compared to baseline. |
| Cost drift | Input / output token ratio shifts for the same prompts. | Token counts per canary run plus any unit-cost change to flag. |
| Lifecycle drift | Whether the catalogue's lifecycle field for the snapshot has shifted. | Lifecycle status + retrievedAt at each canary run; flag any change. |
Record this in your brief
Brief content that keeps the playbook output reusable for a reviewer.
- Embed the canary set + reference rubric so the regression record is self-contained.
- Pair drift observations with the snapshot ID under test.
- List the pass-rate floor and the action taken on breaches.
- Cross-reference any failing canary with /reverification before treating it as a vendor allegation.
Related templates
Related workflows
What the lab does not promise
- No production readiness guarantee. A passing playbook is evidence, not approval.
- No compliance or regulatory certification. Verification is not certification.
- No safety validation. Templates and playbooks are planning tools, not safety reviews.
- No model ranking. The lab does not score candidates against each other.
- No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.