Skip to content
WebmasterID

Lab · intermediate

Model regression testing

How to run a small, repeatable canary suite after every snapshot rotation so silent regressions surface before production traffic notices.

Difficulty

intermediate

Estimated time

30 min

Output

2 Markdown artifacts

A planning recipe, not a safety validation.

The canary suite catches drift; it does not certify the model. A passing canary is not production readiness, and a failing canary is not a vendor allegation — investigate before escalating.

Goal

Catch a silent regression introduced by a snapshot rotation before downstream users hit it.

When to use this

  • Your provider rotates model snapshots without changing the API model name.
  • Your automation runs unattended for hours or days at a time.
  • You have an evidence brief from a previous selection round and want to keep it honest.

Prerequisites

Test setup

  • Freeze a canary suite of 10–25 prompts that exercises the behaviours your application depends on.
  • Store reference outputs from the snapshot you originally selected — keep them out of source control if they contain real data.
  • Schedule the canary suite to run on a regular cadence (every release, every snapshot bump, or every 24 hours).
  • Wire alerting that fires when canary pass rate drops below your floor.

Minimum test set

  • 10–25 representative prompts you also ran in the original selection.
  • At least one prompt per behaviour your application depends on (recall, format, refusal, modality).

Prompt variants

Vary one dimension at a time so you can attribute behaviour.

  • Same prompt, same parameters, two snapshots — reference vs current.
  • Same prompt with and without retrieved context to separate model drift from retrieval drift.

Observations to record

Record observations, not scores. The evidence brief stays auditable.

  • Pass / fail per prompt against reference output (or rubric, if exact match is not appropriate).
  • Latency drift over time.
  • Cost drift over time (input/output token shifts can move cost without changing pricing).
  • Lifecycle field changes against /reverification.

Failure modes to watch

  • Pass rate drops sharply on one prompt category — likely targeted regression.
  • Pass rate drops gradually across categories — likely a generic snapshot drift.
  • Cost climbs while pass rate stays flat — verbose-output drift.
  • The provider deprecates the snapshot mid-run — confirm with the catalogue's lifecycle field.

Stop conditions

Knowing when to stop is part of the test.

  • Stop and re-evaluate if pass rate drops below your floor on any release.
  • Stop and re-verify if the catalogue's lifecycle field shifts to deprecated for the snapshot under test.

Outputs

  • A Markdown evidence delta listing the prompts that regressed.
  • An updated decision brief if the regression triggers a re-selection.

Weak test vs stronger test

Weak test

  • Run a freeform smoke test ad-hoc and call it a regression check.
  • Skip storing reference outputs from the originally-selected snapshot.
  • Treat a single failing canary as a vendor issue without investigating.

Stronger test

  • Freeze a canary suite of 10–25 prompts covering the behaviours the application depends on.
  • Store reference outputs (or rubric criteria) for comparison.
  • Schedule the suite on a regular cadence and wire alerting against a documented pass-rate floor.
  • Investigate a failing canary against /reverification before escalating.

Observation rubric

Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.

DimensionWhat to look forWhat to record
Per-prompt driftWhether each canary prompt still meets the reference expectation.Pass / fail per prompt with rationale, not a rolled-up percentage.
Latency driftWhether latency for the canary set has shifted from the baseline.Median + tail latency per canary run, compared to baseline.
Cost driftInput / output token ratio shifts for the same prompts.Token counts per canary run plus any unit-cost change to flag.
Lifecycle driftWhether the catalogue's lifecycle field for the snapshot has shifted.Lifecycle status + retrievedAt at each canary run; flag any change.

Record this in your brief

Brief content that keeps the playbook output reusable for a reviewer.

  • Embed the canary set + reference rubric so the regression record is self-contained.
  • Pair drift observations with the snapshot ID under test.
  • List the pass-rate floor and the action taken on breaches.
  • Cross-reference any failing canary with /reverification before treating it as a vendor allegation.

Related templates

Related workflows

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.