Skip to content
WebmasterID

Exercise · intermediate

Plan an external model test

Use the brief and the testing lesson to plan your own prompt, latency, rate-limit, and cost validation work for the shortlist.

Estimated time: 12 minutes. Exercises produce evidence artifacts, never model recommendations.

Goal

Leave the catalogue with a written test plan that pairs the brief with the workload-specific tests only your team can run.

Prerequisites

  • Read /learn/testing-ai-models so the prompt/latency/cost/compliance categories are clear.
  • Have a brief Markdown export from the create-decision-brief exercise.

Step-by-step exercise

  1. Re-open your brief

    Re-export the brief from /briefs/build or /api/briefs/decision so you have the current evidence pack.

    Open /briefs/build →

    Expected outcome: Your brief is current as of today's build date.

  2. Read the data gaps section

    The brief's data-gap list is your test plan input. For each gap, write down which test will resolve it (prompt test, latency test, rate-limit test, cost projection, compliance review).

    Open /briefs/build →

    Expected outcome: Every data gap has a named test that will close it.

  3. Add a prompt test

    Choose 3–5 real prompts (or representative ones) and a quality rubric. Note where each shortlisted model will be called from.

    Open /learn/testing-ai-models →

    Expected outcome: You have a prompt test plan with at least one acceptance criterion per prompt.

  4. Add a cost projection

    Multiply your expected traffic mix by each shortlisted model's verified per-unit pricing references. Watch for unit semantics across providers.

    Open /learn/pricing-references →

    Expected outcome: You have a per-model monthly cost projection your finance reviewer can sanity-check.

  5. Pair the brief with your test plan

    Attach the brief Markdown to your test plan as the catalogue's contribution; your tests are your team's contribution.

    Open /examples/decision-brief →

    Expected outcome: Reviewer pack = brief + test plan. Neither one alone is the artifact.

Completion checklist

Completion checklist

  • Every data gap in the brief is paired with a planned test.
  • Prompt test plan has acceptance criteria.
  • Cost projection accounts for unit semantics.
  • You did NOT skip compliance review because the model was 'verified'.

Evidence artifact

At the end of this exercise you should have: A written test plan (Markdown or doc) that pairs your brief with the workload-specific tests.

Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.

Related workflow routes

Exercise does not recommend a model — external testing still required.

Common mistake

Reusing benchmark numbers as the test plan. The plan should name your prompts, your acceptance rubric, your region, and your traffic mix — not republish a vendor figure.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Example artifact

## External test plan (illustrative)
Candidate: <slug> @ <snapshot>
Region: <inference-region>

Tests:
- Prompt set: 5–10 representative prompts with rubric per category
- Latency: measured from <region> at <load>
- Rate limits: deliberate burst at <RPS>
- Cost: project from <prompts/day> × pricing reference
- Compliance: review against <regime / internal control>
- Regression: canary suite scheduled <cadence>

Reviewer sign-offs required: <names / roles>

Substitute your real shortlist, slugs, dates, and values when you walk the exercise.

Repeat this exercise when

  • Before any production launch.
  • After every snapshot rotation.
  • Whenever the brief's data gaps change.

Review before moving on

  • Prompt set + rubric are written down.
  • Latency region matches production.
  • Cost projection uses a fresh pricing reference.
  • Regression cadence is scheduled, not aspirational.
  • Compliance review is named, not skipped.

Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.

What this exercise does not produce

  • A model recommendation. The exercise routes you through evidence — you decide.
  • A score or grade for any model. The catalogue does not score.
  • A substitute for external prompt, latency, rate-limit, cost, or compliance tests.