Skip to content
WebmasterID

Lab · Prompts

Evaluation prompt library

Prompt sets for testing model behaviour before production use. These are evaluation inputs you run in your own model harness, not production prompts. Outputs feed into the prompt-test matrix template and the decision brief builder.

Start here

Prompt sets

6 prompt sets

Each set targets one evaluation dimension. Open a set to see the prompts, expected observations, failure modes, and what to record.

How to use these prompts

  1. Choose a prompt set that matches the behaviour you need to evaluate.
  2. Run the same prompts across candidate models in your own environment — your keys, your region, your sampling parameters.
  3. Record outputs in the prompt-test matrix template.
  4. Compare observations, not vibes. Record per-prompt evidence rather than collapsing to a single score.
  5. Add findings to a decision brief for the next reviewer.

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.