Skip to content
WebmasterID

Learn · testing workflow

How to test an AI model before integration

After the shortlist: how to run your own prompt, latency, rate-limit, cost, and compliance tests — using the evidence brief as the pack you ship to reviewers.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

After the shortlist, you still need to test

The catalogue can take you from "I have a use case" to "I have a shortlist of three or four candidate models with verified fields and a paste-ready evidence brief". What it cannot do is run your workload against those models. Every selection workflow ends at a testing phase the team owns.

The good news is that the evidence brief from /briefs/build is exactly the kind of artifact a reviewer wants to see alongside your test results: a clear record of which fields the catalogue knew, which it did not, and where every claim came from.

What to test, in order

Workload-specific tests the catalogue cannot run

  • Prompt tests — your real prompts (or representative ones), with your real retrieved context. Measure output quality against your own rubric.
  • Output structure — does the model reliably produce the structured output (JSON, tool calls, structured fields) your application needs?
  • Latency in your environment — measured from your region against the chosen endpoint, with realistic prompt sizes.
  • Rate limits — does the provider's published rate limit cover your expected peak? Run a deliberate burst to confirm.
  • Cost validation — multiply your real traffic pattern by the per-unit pricing reference. Watch for prompt-size tiers, cache write/read accounting, and output-token differences.
  • Compliance + security review — what data leaves your environment, where is it processed, what is logged on the provider's side, what retention applies?
  • Failure modes — what does the model do when the prompt is malformed, hostile, or far outside the training distribution?

Use the decision brief as your evidence pack

The decision brief is built to be a reviewer's reading pack: every field is captured with its citation, every data gap is listed explicitly, and the comparison column count is bounded. Attach the brief to whatever testing artifacts your team produces — a test plan, latency measurements, a cost projection — and ship them together. The brief gives the reviewer the catalogue's contribution; the tests give them the workload-specific contribution. Together they answer "why this model" with evidence rather than opinion.

Common testing failure modes

Common mistakes

  • Testing one prompt and generalising.

    Model quality varies wildly across prompt shapes. A single example is anecdote, not evidence.

  • Measuring latency from your laptop on a home network.

    Production latency is a function of your serving region, the model's serving region, and the network path. Test from where your app actually runs.

  • Estimating cost from a single per-token rate.

    Real cost depends on input vs output ratio, cache hit rate, prompt-size tiers, and any retry overhead. Project from your actual traffic mix.

  • Skipping compliance review because the catalogue 'said the model is verified'.

    Verification means the catalogue confirmed a value against a primary source. It does not assert that the model meets any specific regulatory regime — that review is yours.

  • Not retesting after a snapshot rotation.

    Providers update their underlying weights even when the model name stays the same. A passing evaluation last quarter is not a passing evaluation today.

Apply this workflow

Apply this workflow

Practise this lesson

These exercises route the lesson concept through the verified-data product surfaces. Each one ends with a concrete artifact you can share.

Data gaps the catalogue intentionally does not fill

Latency, throughput, uptime, and workload-specific quality measurements are intentionally absent from the catalogue. These are environment-dependent and would mislead more often than help. Your own tests are the only honest source for them. The catalogue's job is to give you the verified catalogue fields and a paste-ready evidence brief — the testing is yours.

Related pages

Sources and freshness

The values in the evidence brief carry their citations forward, so the reviewer reading your test results can trace every catalogue claim back to its source page. Re-export the brief before the review meeting if the catalogue has updated any of the underlying fields — the reverification queue shows what has moved recently.

Teaching example

Illustrative — not a recommendation.

Situation: The catalogue produced a shortlist + brief, but the team is about to integrate without running its own workload-specific tests.

Decision to make: What is the smallest test plan that meaningfully reduces the integration risk before launch?

Verified fields that matter:

  • Prompt set (5–10 representative prompts)
  • Acceptance rubric per prompt category
  • Pinned sampling parameters
  • Pricing reference for the cost projection
  • Lifecycle status (gate the timeline)

Weak vs better approach

Weak approach

  • Run a happy-path prompt once and ship.
  • Estimate cost from a single token count.
  • Measure latency from a developer laptop.
  • Skip the regression test after launch.

Better approach

  • Run 5–10 representative prompts with a pre-agreed acceptance rubric.
  • Project cost from your actual traffic mix.
  • Measure latency from the region the application will serve.
  • Schedule a small canary suite to detect snapshot drift after launch.

Why better: Real workload behaviour shows up only in workload-specific tests. The better approach replaces anecdote with a small, repeatable test plan that survives a snapshot rotation.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

External test plan note

## External test plan (illustrative)
Candidate: <slug> @ <snapshot>
Region: <inference-region>

Test set (5–10 prompts):
- P-01 happy path · rubric: <criteria>
- P-02 edge case · rubric: <criteria>
- P-03 adversarial · rubric: <criteria>

Sampling: temperature <value>, top_p <value>, max_tokens <value>
Cost projection: <prompts/day> × <input+output tokens> × <pricing-ref>
Regression cadence: every release OR weekly canary
Reviewer sign-off required: <names/roles>

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Read the testing framework the catalogue does not run for you.

  2. Step 2

    Apply in /briefs/build →

    Attach the test plan to the evidence brief.

  3. Step 3

    Verify in /sources →

    Re-check the citations the brief depends on before launch.

  4. Step 4

    Test in /lab →

    Walk the minimum prompt-testing routine end to end.

Review before moving on

  • I have 5–10 representative prompts with acceptance criteria.
  • Sampling parameters are pinned and recorded.
  • Cost projection uses a fresh pricing reference.
  • Latency is measured from the inference region we will serve.
  • A regression suite is scheduled after launch.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Doing the testing for you — the catalogue surfaces fields; your workload-specific tests are yours.
  • Asserting which model wins your evaluation — that depends on your data, traffic, and constraints.
  • Certifying the model for your compliance regime — verification ≠ certification.