Skip to content
WebmasterID

Lab · intermediate

Long-context testing

How to test long-prompt behaviour past the catalogue's verified context window — recall, instruction adherence, and cost growth — without trusting a marketing number.

Difficulty

intermediate

Estimated time

35 min

Output

2 Markdown artifacts

A planning recipe, not a safety validation.

Long-context behaviour is workload-specific. The playbook does not assert a model's effective context length and does not rank candidates by long-prompt cost.

Goal

Decide whether a candidate model holds up at the prompt size your real workload sends.

When to use this

  • Your prompts plus retrieved context plus expected output approach the verified context window.
  • You need to understand cost growth at large prompt sizes before committing.
  • You suspect the model's effective context is smaller than its advertised context.

Prerequisites

Test setup

  • Build a prompt scaffold that lets you append filler tokens to grow the prompt without changing the question.
  • Pin a deterministic token counter so prompt sizes are comparable across runs.
  • Capture per-call latency separately from any RAG retrieval latency.
  • Have the pricing reference and unit semantics open so cost projections stay honest.

Minimum test set

  • The same question asked at 4k / 32k / 128k / 256k+ token prompt sizes (cap at the verified context window).
  • A recall-style prompt where the answer is buried at start, middle, and end of the input.
  • An instruction-adherence prompt where the system prompt and the buried instruction disagree.

Prompt variants

Vary one dimension at a time so you can attribute behaviour.

  • Filler placement — same total tokens, but the answer sits at start, middle, or end.
  • Retrieval shuffling — same chunks, different order, to test order sensitivity.
  • Compression — the same answer with and without irrelevant context.

Observations to record

Record observations, not scores. The evidence brief stays auditable.

  • Pass / fail per recall position.
  • Latency growth as prompt size grows.
  • Cost growth as prompt size grows (input vs output tokens, separately).
  • Whether output structure degrades as prompt size grows.

Failure modes to watch

  • Model passes the recall test at start but degrades severely mid-prompt.
  • Cost scales superlinearly because the provider tiers pricing on prompt length.
  • Latency hits a hard ceiling above a certain prompt size.
  • The model truncates output silently when the input nears the verified context limit.

Stop conditions

Knowing when to stop is part of the test.

  • Stop and re-scope if the model degrades below your acceptance rate at prompt sizes you actually need.
  • Stop and document if cost growth violates your finance ceiling — that is selection evidence.

Outputs

  • A Markdown evidence brief listing recall and cost behaviour per prompt size per candidate.
  • A short cost projection note attached to /briefs/build for reviewer pickup.

Weak test vs stronger test

Weak test

  • Fill the prompt to the verified context window with random tokens and assume recall stays constant.
  • Skip varying the answer position inside the prompt.
  • Ignore cost growth at large prompt sizes.

Stronger test

  • Build a scaffold that grows the prompt while pinning the question.
  • Test recall with the answer at the start, middle, and end of the input.
  • Capture cost behaviour separately for input vs output tokens.
  • Cap the test at the verified context window, not the marketing one.

Observation rubric

Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.

DimensionWhat to look forWhat to record
Recall by positionWhether the model surfaces the buried answer regardless of its position in the input.Pass / fail per position bucket, with the prompt scaffold captured.
Latency growthLatency as prompt size grows.Latency observation per prompt-size bucket, with the region you ran from.
Cost growthCost growth as prompt size grows (input vs output tokens, separately).Per-call token counts mapped to the pricing reference; flag any tier change.
Output structure degradationWhether structured-output reliability or instruction following degrades at large prompt sizes.Per-bucket failure modes with example outputs.

Record this in your brief

Brief content that keeps the playbook output reusable for a reviewer.

  • Embed the scaffold and the answer-position protocol so the brief is reproducible.
  • Record recall by prompt-size bucket rather than a single rolled-up figure.
  • Pair cost growth observations with the pricing reference + retrievedAt date.
  • Note any prompt size where the model truncated output silently.

Related templates

Related workflows

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.