Skip to content
WebmasterID

Lab · beginner

Prompt testing basics

The minimum prompt-testing routine to run against a shortlisted model before integration. Defines a representative prompt set, structured observations, and concrete failure modes — no benchmark scores.

Difficulty

beginner

Estimated time

25 min

Output

2 Markdown artifacts

A planning recipe, not a safety validation.

The playbook teaches you to test. It does not score the model for you, does not certify the model for any regulatory regime, and does not declare a winner.

Goal

Decide whether a shortlisted model passes your own prompt rubric before you wire it into anything.

When to use this

  • You have 1–4 candidate models from the selection workspace and need to compare them on your own prompts.
  • You want a repeatable testing routine that does not depend on published benchmark scores.
  • You need an evidence trail your reviewer can read independently.

Prerequisites

Test setup

  • Open the candidate models in tabs and confirm each one is in an active lifecycle state.
  • Pick the inference region you will use in production — run tests from that region whenever possible.
  • Set a fixed system prompt for the suite so prompt variance does not contaminate model variance.
  • Decide ahead of time which sampling parameters (temperature, top_p, max_tokens) you will hold constant.

Minimum test set

  • 5–10 representative prompts drawn from your real workload (or close stand-ins).
  • At least one happy-path prompt, one edge-case prompt, and one adversarial prompt per category you ship.
  • A short rubric per prompt that names the acceptance criteria — not a numeric score.

Prompt variants

Vary one dimension at a time so you can attribute behaviour.

  • Same prompt, three temperatures (for example 0.0, 0.3, 0.7) to see how the model behaves under sampling.
  • Same prompt, varied system prompt length to surface instruction-following degradation.
  • Same prompt, with and without retrieved context to see how the model handles RAG noise.

Observations to record

Record observations, not scores. The evidence brief stays auditable.

  • Pass / fail against each acceptance criterion, with a short rationale.
  • Wall-clock latency from your environment for each call (note the region you ran from).
  • Input + output token counts so you can project cost from the pricing reference.
  • Refusal rate and any structured-output validity failures.

Failure modes to watch

  • Model passes the happy path but fails on adversarial input — typical when the candidate model lags on safety training.
  • Model passes single-shot but degrades when you chain prompts.
  • Structured output is valid JSON but does not match your schema constraints.
  • Latency spikes mid-suite — confirm whether it is the model, the region, or your retry strategy.

Stop conditions

Knowing when to stop is part of the test.

  • Stop and re-scope if a candidate fails on more than one happy-path prompt — the catalogue's verified fields say nothing about prompt-level reliability.
  • Stop and request reverification if the model's lifecycle field shifts to deprecated during the run.
  • Stop and document if you cannot replicate a result twice in a row — non-determinism is itself evidence.

Outputs

  • A Markdown evidence brief listing acceptance results per prompt per model.
  • A short note attached to your /briefs/build export naming which prompts you actually ran.

Weak test vs stronger test

Weak test

  • Run one happy-path prompt at default temperature, eyeball the output, and ship.
  • Skip recording outputs verbatim.
  • Pretend a single positive result generalises across prompt categories.

Stronger test

  • Run 5–10 representative prompts spanning happy / edge / adversarial / refusal categories.
  • Hold sampling parameters constant and record the values.
  • Capture every output verbatim before applying acceptance criteria.
  • Note non-determinism across reruns instead of hiding it.

Observation rubric

Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.

DimensionWhat to look forWhat to record
Acceptance against rubricDoes the output meet the pre-agreed acceptance criteria for the prompt category?Pass / fail per criterion, with a short rationale — not a numeric score.
LatencyWall-clock latency from your environment for each call.Median and tail latency per prompt, with the region you ran from.
Token usageInput vs output token counts per call.Counts per prompt, so the cost projection later can use the pricing reference.
Refusal / structured failureRefusal rate and any structured-output validity failures.Counts per category plus a short note on why each failure happened.

Record this in your brief

Brief content that keeps the playbook output reusable for a reviewer.

  • Embed the prompt set + acceptance criteria so the brief is self-contained.
  • Record per-prompt outcomes instead of a single rolled-up percentage.
  • Attach the verbatim outputs for failures so the reviewer can replay them.
  • Note any prompts that triggered the stop condition and why.

Related templates

Related workflows

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.