Lab · Prompts
Evaluation prompt library
Prompt sets for testing model behaviour before production use. These are evaluation inputs you run in your own model harness, not production prompts. Outputs feed into the prompt-test matrix template and the decision brief builder.
Start here
Prompt sets
6 prompt sets
Each set targets one evaluation dimension. Open a set to see the prompts, expected observations, failure modes, and what to record.
- beginner20 min· 5 prompts
summarization
Summarization quality
Evaluate whether a model summarises without adding unsupported claims, omitting constraints, or inventing numbers.
Open prompt set →
- intermediate25 min· 5 prompts
structured-output
Structured extraction
Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.
Open prompt set →
- intermediate25 min· 5 prompts
long-context
Long-context recall
Evaluate whether a model preserves constraints, handles cross-references, and detects conflicts across multiple sections of input.
Open prompt set →
- beginner20 min· 5 prompts
instruction-following
Instruction following
Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.
Open prompt set →
- intermediate25 min· 5 prompts
safety-boundary
Refusal boundary
Evaluate whether the model handles benign boundary-setting safely — without over-refusing, without giving definitive professional advice, and without complying with inappropriate requests.
Open prompt set →
- intermediate25 min· 5 prompts
automation
Automation robustness
Evaluate whether a model handles automation-style constraints (allowed categories, missing values, retry decisions, ambiguity flags) without silently breaking the contract.
Open prompt set →
How to use these prompts
- Choose a prompt set that matches the behaviour you need to evaluate.
- Run the same prompts across candidate models in your own environment — your keys, your region, your sampling parameters.
- Record outputs in the prompt-test matrix template.
- Compare observations, not vibes. Record per-prompt evidence rather than collapsing to a single score.
- Add findings to a decision brief for the next reviewer.
Prompt library policy
- These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
- No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
- No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
- No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
- No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
- No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.