Skip to content
WebmasterID

Lab · long-context

Long-context recall

Evaluate whether a model preserves constraints, handles cross-references, and detects conflicts across multiple sections of input.

Category

long-context

Difficulty

intermediate

Estimated time

25 min

Prompts

5

When to use this

  • Your workload sends multi-section prompts (retrieved context + system + user).
  • You need to confirm a candidate tracks earlier constraints later in the input.
  • You want to surface 'lost in the middle' behaviour without building a full benchmark.

Evaluation goal

Confirm the candidate preserves earlier constraints when answering later questions, flags conflicts honestly, and refuses to fabricate cross-references that are not in the input.

Prerequisites

Export

Open the raw Markdown if you want to paste the whole set into your harness or design doc.

Open raw Markdown → /api/lab/prompts/long-context-recall

Prompts

5 evaluation prompts

Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.

  1. Recall earlier constraint

    Purpose

    Confirm the model recalls a specific constraint from the earliest section.

    Prompt

    Read all three sections, then answer the question.
    
    Section A — Policy: All inference requests above 32k tokens are routed to the long-context pool. The long-context pool publishes pricing per million input tokens separately from per million output tokens. No verified latency claim is published.
    
    Section B — Pricing reference: The catalogue records prompt-size pricing tiers when the provider publishes them. Pricing reference is a verified-field reference, not a live quote.
    
    Section C — Conflicting detail: Earlier internal notes mentioned "all requests over 16k tokens are long-context" — this is NOT current policy. Current policy is the 32k threshold from Section A.
    
    QUESTION: What is the current long-context threshold, according to Section A?

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Answer: 32k tokens.

    Failure looks like

    • Answers 16k tokens (Section C historical detail).
    • Says the threshold is unspecified.
    • Invents a different number.

    What to record

    • Numeric accuracy.
    • Whether the model cites Section A explicitly.
  2. Detect conflicting detail

    Purpose

    Confirm the model surfaces internal conflicts rather than smoothing them over.

    Prompt

    Using the three sections above, list any conflicting details between them. Quote each conflict in the order it appears.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Identifies the 16k vs 32k threshold conflict between Section A and Section C. May note Section C explicitly says the 16k figure is not current policy.

    Failure looks like

    • Reports no conflicts.
    • Reports a conflict that is not in the source.
    • Smooths the conflict by averaging the values.

    What to record

    • Conflict identified vs missed.
    • Any fabricated conflict.
  3. Cross-reference policy + pricing

    Purpose

    Confirm the model handles cross-reference without inventing a relationship.

    Prompt

    Using the three sections above, explain how the long-context routing in Section A relates to the pricing reference in Section B. If the sections do not state the relationship, say so explicitly.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Notes that Section A defines routing and Section B defines how pricing is recorded; explicitly says the sections do not assert a specific pricing tier number.

    Failure looks like

    • Invents a specific tier price.
    • Asserts a cost relationship the sources did not state.
    • Claims Section B contradicts Section A.

    What to record

    • Whether the explicit 'do not state' answer was given when appropriate.
    • Any invented numeric relationship.
  4. Absent-information request

    Purpose

    Confirm the model refuses to invent missing metrics.

    Prompt

    Using the three sections above, state the published verified latency for long-context requests. If the sources do not state a verified latency, reply "Not stated in source." and nothing else.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Answer: "Not stated in source."

    Failure looks like

    • Invents a latency number.
    • Returns a generic latency claim.
    • Returns more than the required short refusal.

    What to record

    • Exact match against the required refusal phrasing.
    • Any invented latency value.
  5. Constraint preservation across order

    Purpose

    Confirm the model preserves an ordering constraint and assigns sections correctly.

    Prompt

    Re-read the sections. Answer in order: (1) which threshold is current policy, (2) which threshold is historical, (3) which section is the load-bearing source for current policy.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    (1) 32k tokens. (2) 16k tokens. (3) Section A.

    Failure looks like

    • Reverses current vs historical.
    • Attributes current policy to Section C.
    • Skips one of the three answers.

    What to record

    • Per-answer correctness.
    • Whether the order was preserved.

Observation checklist

Tick items off in your own notes; the page does not store progress.

  • Did the model attribute current policy correctly to Section A?
  • Did the model flag the 16k vs 32k conflict?
  • Did the model refuse to invent latency or pricing numbers?
  • Did the model preserve ordering when asked for ordered answers?
  • Did the model produce a short, exact refusal when asked?

Comparison notes

Apply when running the set across multiple candidate models.

  • If a candidate produces inconsistent answers across reruns, record the variance — that is observability data, not failure data.
  • Compare candidates on the same section ordering. Shuffling sections is a separate test.
  • Capture the model's section attributions — they are useful evidence even when the final answer is right.

How to use with the prompt-test matrix

The matrix template ships at /lab/templates/prompt-test-matrix.

  • Map each prompt ID (LCR-01 … LCR-05) to a row in the matrix.
  • Record verbatim answers — quote attribution is part of the evidence.
  • Capture section-attribution observations alongside the final answer.

What not to conclude

  • That the model's recall behaviour scales to your real prompt sizes.
  • That absent-information handling generalises beyond this set.
  • That cross-reference reliability holds for asymmetric retrieval orderings.

When to rerun this set

  • Your prompt structure changes (number of sections, ordering).
  • The provider expands the verified context window.
  • A snapshot rotation shifts long-prompt behaviour.

Related playbooks and templates

Related workflows

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.