Lab · beginner
Prompt testing basics
The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.
Difficulty
beginner
Estimated time
25 min
Output
2 Markdown artifacts
A planning recipe, not a safety validation.
The playbook teaches you to test. It does not score the model for you, does not certify the model for any regulatory regime, and does not declare a winner.
Goal
Decide whether a shortlisted model passes your own prompt rubric before you wire it into anything.
When to use this
- You have 1–4 candidate models from the selection workspace and need to compare them on your own prompts.
- You want a repeatable testing routine that does not depend on published benchmark scores.
- You need an evidence trail your reviewer can read independently.
Prerequisites
Test setup
- Open the candidate models in tabs and confirm each one is in an active lifecycle state.
- Pick the inference region you will use in production — run tests from that region whenever possible.
- Set a fixed system prompt for the suite so prompt variance does not contaminate model variance.
- Decide ahead of time which sampling parameters (temperature, top_p, max_tokens) you will hold constant.
Minimum test set
- 5–10 representative prompts drawn from your real workload (or close stand-ins).
- At least one happy-path prompt, one edge-case prompt, and one adversarial prompt per category you ship.
- A short rubric per prompt that names the acceptance criteria — not a numeric score.
Prompt variants
Vary one dimension at a time so you can attribute behaviour.
- Same prompt, three temperatures (for example 0.0, 0.3, 0.7) to see how the model behaves under sampling.
- Same prompt, varied system prompt length to surface instruction-following degradation.
- Same prompt, with and without retrieved context to see how the model handles RAG noise.
Observations to record
Record observations, not scores. The evidence brief stays auditable.
- Pass / fail against each acceptance criterion, with a short rationale.
- Wall-clock latency from your environment for each call (note the region you ran from).
- Input + output token counts so you can project cost from the pricing reference.
- Refusal rate and any structured-output validity failures.
Failure modes to watch
- Model passes the happy path but fails on adversarial input — typical when the candidate model lags on safety training.
- Model passes single-shot but degrades when you chain prompts.
- Structured output is valid JSON but does not match your schema constraints.
- Latency spikes mid-suite — confirm whether it is the model, the region, or your retry strategy.
Stop conditions
Knowing when to stop is part of the test.
- Stop and re-scope if a candidate fails on more than one happy-path prompt — the catalogue's verified fields say nothing about prompt-level reliability.
- Stop and request reverification if the model's lifecycle field shifts to deprecated during the run.
- Stop and document if you cannot replicate a result twice in a row — non-determinism is itself evidence.
Outputs
- A Markdown evidence brief listing acceptance results per prompt per model.
- A short note attached to your /briefs/build export naming which prompts you actually ran.
Weak test vs stronger test
Weak test
- Run one happy-path prompt at default temperature, eyeball the output, and ship.
- Skip recording outputs verbatim.
- Pretend a single positive result generalises across prompt categories.
Stronger test
- Run 5–10 representative prompts spanning happy / edge / adversarial / refusal categories.
- Hold sampling parameters constant and record the values.
- Capture every output verbatim before applying acceptance criteria.
- Note non-determinism across reruns instead of hiding it.
Observation rubric
Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.
| Dimension | What to look for | What to record |
|---|---|---|
| Acceptance against rubric | Does the output meet the pre-agreed acceptance criteria for the prompt category? | Pass / fail per criterion, with a short rationale — not a numeric score. |
| Latency | Wall-clock latency from your environment for each call. | Median and tail latency per prompt, with the region you ran from. |
| Token usage | Input vs output token counts per call. | Counts per prompt, so the cost projection later can use the pricing reference. |
| Refusal / structured failure | Refusal rate and any structured-output validity failures. | Counts per category plus a short note on why each failure happened. |
Record this in your brief
Brief content that keeps the playbook output reusable for a reviewer.
- Embed the prompt set + acceptance criteria so the brief is self-contained.
- Record per-prompt outcomes instead of a single rolled-up percentage.
- Attach the verbatim outputs for failures so the reviewer can replay them.
- Note any prompts that triggered the stop condition and why.
Related templates
Related workflows
What the lab does not promise
- No production readiness guarantee. A passing playbook is evidence, not approval.
- No compliance or regulatory certification. Verification is not certification.
- No safety validation. Templates and playbooks are planning tools, not safety reviews.
- No model ranking. The lab does not score candidates against each other.
- No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.