Skip to content
WebmasterID

Lab · instruction-following

Instruction following

Evaluate whether a model honours formatting, word-count, uncertainty, and forbidden-phrase instructions without silent drift.

Category

instruction-following

Difficulty

beginner

Estimated time

20 min

Prompts

5

When to use this

  • You depend on the model returning output in a specific format.
  • You need to confirm the model expresses uncertainty rather than fabricating confidence.
  • You want a small set you can rerun after a snapshot rotation.

Evaluation goal

Confirm the candidate follows explicit format, length, uncertainty, and forbidden-phrase instructions consistently.

Prerequisites

Export

Open the raw Markdown if you want to paste the whole set into your harness or design doc.

Open raw Markdown → /api/lab/prompts/instruction-following

Prompts

5 evaluation prompts

Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.

  1. Exact format

    Purpose

    Confirm the model returns the exact requested format.

    Prompt

    Reply with exactly three lines:
    Line 1: TOPIC = catalogue freshness
    Line 2: STATUS = stable
    Line 3: NEXT_STEP = re-read /sources weekly
    Do not add a fourth line. Do not add explanatory prose.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Three lines, in the exact order requested, with no extra commentary.

    Failure looks like

    • Adds a fourth line.
    • Adds prose before or after the lines.
    • Reorders the three lines.

    What to record

    • Line count.
    • Any added commentary.
    • Order of the three lines.
  2. Word limit

    Purpose

    Confirm the model honours a word-count cap.

    Prompt

    Explain in 25 words or fewer why a pricing reference is not the same as a live quote.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response is 25 words or fewer and addresses the difference between a reference and a live quote.

    Failure looks like

    • Exceeds 25 words.
    • Does not address the difference.
    • Adds a separate disclaimer that pushes the response over the cap.

    What to record

    • Word count.
    • Whether the answer addresses the prompt.
  3. No unsupported claim

    Purpose

    Confirm the model honours a no-specific-claim constraint.

    Prompt

    In 1–3 sentences, describe how the verified-field discipline works on a catalogue page. Do not name any specific model, provider, or dollar amount. If you cannot answer without naming one, reply "Cannot answer under the constraint."

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response stays generic — no specific model name, no specific provider, no dollar amount — or returns the exact fallback string.

    Failure looks like

    • Names a specific model or provider.
    • Quotes a dollar amount.
    • Returns an empty response.

    What to record

    • Whether constraint was honoured.
    • Any specific names or amounts that leaked.
  4. Uncertainty expression

    Purpose

    Confirm the model expresses uncertainty when it has no grounding.

    Prompt

    If you are not certain of the answer to the following question, say so explicitly with the phrase "I am not certain because: ..." followed by the reason. Question: Has the catalogue's pricing reference for the Atlas provider been re-verified in the past 7 days?

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response begins with "I am not certain because:" and gives a plausible reason (no access to a live source, no provided context, etc.).

    Failure looks like

    • Asserts a definitive yes or no.
    • Skips the required phrase.
    • Fabricates a verification timestamp.

    What to record

    • Did the response begin with the required phrase?
    • Did the model assert certainty it could not have?
  5. Forbidden phrase avoidance

    Purpose

    Confirm the model can avoid a specified vocabulary.

    Prompt

    In 2–3 sentences, describe what /coverage shows. Do not use the words "best", "winner", "guaranteed", or "certified" anywhere in your response.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response avoids all four forbidden words while still answering the question.

    Failure looks like

    • Includes any of the forbidden words.
    • Returns a refusal even though the request is benign.
    • Returns a single forbidden-word synonym that obviously violates the spirit of the constraint.

    What to record

    • Per-word violations.
    • Whether the answer addressed the prompt.

Observation checklist

Tick items off in your own notes; the page does not store progress.

  • Did the model honour exact-format instructions?
  • Did the model respect the word-count cap?
  • Did the model express uncertainty when prompted?
  • Did the model avoid forbidden vocabulary?
  • Did the model silently drop an instruction when it conflicted with its default style?

Comparison notes

Apply when running the set across multiple candidate models.

  • Instruction-following often degrades at higher temperatures — record temperature alongside outcomes.
  • A model that follows instructions once may not follow them on a re-run; record per-prompt re-run consistency.
  • Do not collapse the suite into a percentage — record per-prompt outcomes.

How to use with the prompt-test matrix

The matrix template ships at /lab/templates/prompt-test-matrix.

  • Map each prompt ID (IF-01 … IF-05) to a row.
  • Record per-instruction observations — line count, word count, forbidden-word violations, exact phrase honoured.
  • Note sampling parameters; instruction following often drifts at higher temperatures.

What not to conclude

  • That the model 'follows instructions reliably' across all prompt shapes.
  • That a single passing rerun rules out non-determinism.
  • That instruction adherence in this set predicts production system-prompt behaviour.

When to rerun this set

  • You change the production system prompt shape.
  • Sampling parameters shift in production.
  • A snapshot rotation lands.

Related playbooks and templates

Related workflows

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.