What evaluation means here
Evaluation in this lab is a small, repeatable observation routine. You pick a behaviour you need to validate (faithful summarisation, schema-conformant extraction, long-context recall, instruction following, safe refusal, automation contract adherence). You run a fixed set of prompts across the candidate models. You record outputs and per-prompt acceptance against a rubric you decided in advance. You attach the record to a decision brief for the next reviewer.
Evaluation is not a leaderboard, not a percentage score, and not a vendor pitch. It is the catalogue's contribution to a decision the reader's team owns.
Why prompt sets are not benchmarks
Published benchmarks have specific methodologies, scoring rubrics, and reproducibility properties. The catalogue's evaluation prompt sets share none of those properties — they are small (5 prompts per set), workload-agnostic, and designed to surface a particular failure mode rather than to produce a comparable number.
Treat a prompt set's pass count as evidence for one moment in time, with one set of sampling parameters, against one snapshot of one model. Do not extrapolate to "the candidate is X% better than the baseline."
How to run candidate model trials
- Pick the candidates from a shortlist URL produced in /select.
- Pin sampling parameters (temperature, top_p, max tokens) and hold them constant across the suite.
- Pick one prompt set from /lab/prompts and one playbook from /lab.
- Run the prompts in your own model harness. The catalogue never calls a live model on your behalf.
- Record outputs verbatim. Capture raw responses before any validation step so failures can be replayed.
- Roll observations into the prompt-test matrix template.
- Attach the matrix + the brief to your reviewer pack.
How to record observations
Record what the model did, not what you think the model did. Each prompt has an expected observation, a failure-looks-like list, and a what-to-record list. Stick to those fields. Avoid collapsing observations into "good" or "bad" — the reviewer needs the verbatim output, the acceptance flag, and a short rationale.
If a candidate produces inconsistent answers across reruns, record the variance. Non-determinism is itself evidence.
How to avoid overclaiming
- Do not write "this model is better at structured output" from a 5-prompt run. Write "this model produced schema-valid JSON on 4 of 5 prompts at temperature 0.0 on snapshot X."
- Do not infer latency or cost from a handful of calls. Latency observations require a measurement plan; the catalogue does not publish either.
- Do not promote an evaluation prompt into a production prompt. Evaluation inputs and production inputs are shaped differently.
- Do not extend the suite mid-run. If you need a new prompt, add it to a new run and re-record everything.
How to feed results into a decision brief
Open /briefs/build and select the candidate models you ran the evaluation against. The brief renders verified catalogue fields side by side; you add the evaluation observations underneath. The brief is the artifact your reviewer reads — not the raw matrix.
If the evaluation results contradict the catalogue's verified fields, do not silently reconcile. Record both, attach the source citation, and flag the contradiction in the brief so the reviewer can investigate.