Goal
Leave the catalogue with a written test plan that pairs the brief with the workload-specific tests only your team can run.
Prerequisites
- Read /learn/testing-ai-models so the prompt/latency/cost/compliance categories are clear.
- Have a brief Markdown export from the create-decision-brief exercise.
Step-by-step exercise
Re-open your brief
Re-export the brief from /briefs/build or /api/briefs/decision so you have the current evidence pack.
Expected outcome: Your brief is current as of today's build date.
Read the data gaps section
The brief's data-gap list is your test plan input. For each gap, write down which test will resolve it (prompt test, latency test, rate-limit test, cost projection, compliance review).
Expected outcome: Every data gap has a named test that will close it.
Add a prompt test
Choose 3–5 real prompts (or representative ones) and a quality rubric. Note where each shortlisted model will be called from.
Open
/learn/testing-ai-models→Expected outcome: You have a prompt test plan with at least one acceptance criterion per prompt.
Add a cost projection
Multiply your expected traffic mix by each shortlisted model's verified per-unit pricing references. Watch for unit semantics across providers.
Open
/learn/pricing-references→Expected outcome: You have a per-model monthly cost projection your finance reviewer can sanity-check.
Pair the brief with your test plan
Attach the brief Markdown to your test plan as the catalogue's contribution; your tests are your team's contribution.
Open
/examples/decision-brief→Expected outcome: Reviewer pack = brief + test plan. Neither one alone is the artifact.
Completion checklist
Completion checklist
- Every data gap in the brief is paired with a planned test.
- Prompt test plan has acceptance criteria.
- Cost projection accounts for unit semantics.
- You did NOT skip compliance review because the model was 'verified'.
Evidence artifact
At the end of this exercise you should have: A written test plan (Markdown or doc) that pairs your brief with the workload-specific tests.
Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.
Related workflow routes
- /select — narrow the source-backed shortlist.
- /compare/build — render verified fields side by side.
- /briefs/build — generate the evidence decision brief.
- /sources — every primary-source citation, by provider.
- /coverage — per-provider verified-field coverage.
- /reverification — sources due for manual re-check.
Exercise does not recommend a model — external testing still required.
Common mistake
Reusing benchmark numbers as the test plan. The plan should name your prompts, your acceptance rubric, your region, and your traffic mix — not republish a vendor figure.
Example artifact
Illustrative example — not a recommendation. Substitute your own values when you run the workflow.
Example artifact
## External test plan (illustrative)
Candidate: <slug> @ <snapshot>
Region: <inference-region>
Tests:
- Prompt set: 5–10 representative prompts with rubric per category
- Latency: measured from <region> at <load>
- Rate limits: deliberate burst at <RPS>
- Cost: project from <prompts/day> × pricing reference
- Compliance: review against <regime / internal control>
- Regression: canary suite scheduled <cadence>
Reviewer sign-offs required: <names / roles>Substitute your real shortlist, slugs, dates, and values when you walk the exercise.
Repeat this exercise when
- Before any production launch.
- After every snapshot rotation.
- Whenever the brief's data gaps change.
Review before moving on
- Prompt set + rubric are written down.
- Latency region matches production.
- Cost projection uses a fresh pricing reference.
- Regression cadence is scheduled, not aspirational.
- Compliance review is named, not skipped.
Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.
What this exercise does not produce
- A model recommendation. The exercise routes you through evidence — you decide.
- A score or grade for any model. The catalogue does not score.
- A substitute for external prompt, latency, rate-limit, cost, or compliance tests.