Skip to content
WebmasterID

Lab · Evaluation guide

Evaluation guide

How playbooks, templates, and evaluation prompt sets fit together. The guide walks you through what evaluation means here, why prompt sets are not benchmarks, how to run candidate trials, and how to keep observations honest.

What evaluation means here

Evaluation in this lab is a small, repeatable observation routine. You pick a behaviour you need to validate (faithful summarisation, schema-conformant extraction, long-context recall, instruction following, safe refusal, automation contract adherence). You run a fixed set of prompts across the candidate models. You record outputs and per-prompt acceptance against a rubric you decided in advance. You attach the record to a decision brief for the next reviewer.

Evaluation is not a leaderboard, not a percentage score, and not a vendor pitch. It is the catalogue's contribution to a decision the reader's team owns.

Why prompt sets are not benchmarks

Published benchmarks have specific methodologies, scoring rubrics, and reproducibility properties. The catalogue's evaluation prompt sets share none of those properties — they are small (5 prompts per set), workload-agnostic, and designed to surface a particular failure mode rather than to produce a comparable number.

Treat a prompt set's pass count as evidence for one moment in time, with one set of sampling parameters, against one snapshot of one model. Do not extrapolate to "the candidate is X% better than the baseline."

How to run candidate model trials

  1. Pick the candidates from a shortlist URL produced in /select.
  2. Pin sampling parameters (temperature, top_p, max tokens) and hold them constant across the suite.
  3. Pick one prompt set from /lab/prompts and one playbook from /lab.
  4. Run the prompts in your own model harness. The catalogue never calls a live model on your behalf.
  5. Record outputs verbatim. Capture raw responses before any validation step so failures can be replayed.
  6. Roll observations into the prompt-test matrix template.
  7. Attach the matrix + the brief to your reviewer pack.

How to record observations

Record what the model did, not what you think the model did. Each prompt has an expected observation, a failure-looks-like list, and a what-to-record list. Stick to those fields. Avoid collapsing observations into "good" or "bad" — the reviewer needs the verbatim output, the acceptance flag, and a short rationale.

If a candidate produces inconsistent answers across reruns, record the variance. Non-determinism is itself evidence.

How to avoid overclaiming

  • Do not write "this model is better at structured output" from a 5-prompt run. Write "this model produced schema-valid JSON on 4 of 5 prompts at temperature 0.0 on snapshot X."
  • Do not infer latency or cost from a handful of calls. Latency observations require a measurement plan; the catalogue does not publish either.
  • Do not promote an evaluation prompt into a production prompt. Evaluation inputs and production inputs are shaped differently.
  • Do not extend the suite mid-run. If you need a new prompt, add it to a new run and re-record everything.

How to feed results into a decision brief

Open /briefs/build and select the candidate models you ran the evaluation against. The brief renders verified catalogue fields side by side; you add the evaluation observations underneath. The brief is the artifact your reviewer reads — not the raw matrix.

If the evaluation results contradict the catalogue's verified fields, do not silently reconcile. Record both, attach the source citation, and flag the contradiction in the brief so the reviewer can investigate.

Suggested order through the lab

  1. Step 1

    Read prompt testing basics →
  2. Step 2

    Pick a prompt set →
  3. Step 3

    Open the prompt test matrix →
  4. Step 4

    Run trials in your harness →
  5. Step 5

    Generate a decision brief →

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.