Lab · intermediate
Long-context testing
How to test long-prompt behaviour past the catalogue's verified context window — recall, instruction adherence, and cost growth — without trusting a marketing number.
Difficulty
intermediate
Estimated time
35 min
Output
2 Markdown artifacts
A planning recipe, not a safety validation.
Long-context behaviour is workload-specific. The playbook does not assert a model's effective context length and does not rank candidates by long-prompt cost.
Goal
Decide whether a candidate model holds up at the prompt size your real workload sends.
When to use this
- Your prompts plus retrieved context plus expected output approach the verified context window.
- You need to understand cost growth at large prompt sizes before committing.
- You suspect the model's effective context is smaller than its advertised context.
Prerequisites
Test setup
- Build a prompt scaffold that lets you append filler tokens to grow the prompt without changing the question.
- Pin a deterministic token counter so prompt sizes are comparable across runs.
- Capture per-call latency separately from any RAG retrieval latency.
- Have the pricing reference and unit semantics open so cost projections stay honest.
Minimum test set
- The same question asked at 4k / 32k / 128k / 256k+ token prompt sizes (cap at the verified context window).
- A recall-style prompt where the answer is buried at start, middle, and end of the input.
- An instruction-adherence prompt where the system prompt and the buried instruction disagree.
Prompt variants
Vary one dimension at a time so you can attribute behaviour.
- Filler placement — same total tokens, but the answer sits at start, middle, or end.
- Retrieval shuffling — same chunks, different order, to test order sensitivity.
- Compression — the same answer with and without irrelevant context.
Observations to record
Record observations, not scores. The evidence brief stays auditable.
- Pass / fail per recall position.
- Latency growth as prompt size grows.
- Cost growth as prompt size grows (input vs output tokens, separately).
- Whether output structure degrades as prompt size grows.
Failure modes to watch
- Model passes the recall test at start but degrades severely mid-prompt.
- Cost scales superlinearly because the provider tiers pricing on prompt length.
- Latency hits a hard ceiling above a certain prompt size.
- The model truncates output silently when the input nears the verified context limit.
Stop conditions
Knowing when to stop is part of the test.
- Stop and re-scope if the model degrades below your acceptance rate at prompt sizes you actually need.
- Stop and document if cost growth violates your finance ceiling — that is selection evidence.
Outputs
- A Markdown evidence brief listing recall and cost behaviour per prompt size per candidate.
- A short cost projection note attached to /briefs/build for reviewer pickup.
Weak test vs stronger test
Weak test
- Fill the prompt to the verified context window with random tokens and assume recall stays constant.
- Skip varying the answer position inside the prompt.
- Ignore cost growth at large prompt sizes.
Stronger test
- Build a scaffold that grows the prompt while pinning the question.
- Test recall with the answer at the start, middle, and end of the input.
- Capture cost behaviour separately for input vs output tokens.
- Cap the test at the verified context window, not the marketing one.
Observation rubric
Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.
| Dimension | What to look for | What to record |
|---|---|---|
| Recall by position | Whether the model surfaces the buried answer regardless of its position in the input. | Pass / fail per position bucket, with the prompt scaffold captured. |
| Latency growth | Latency as prompt size grows. | Latency observation per prompt-size bucket, with the region you ran from. |
| Cost growth | Cost growth as prompt size grows (input vs output tokens, separately). | Per-call token counts mapped to the pricing reference; flag any tier change. |
| Output structure degradation | Whether structured-output reliability or instruction following degrades at large prompt sizes. | Per-bucket failure modes with example outputs. |
Record this in your brief
Brief content that keeps the playbook output reusable for a reviewer.
- Embed the scaffold and the answer-position protocol so the brief is reproducible.
- Record recall by prompt-size bucket rather than a single rolled-up figure.
- Pair cost growth observations with the pricing reference + retrievedAt date.
- Note any prompt size where the model truncated output silently.
Related templates
Related workflows
- Long-context demo
- Decision brief builder
- Create a decision brief after testing
- Review sources and freshness before signing off.
What the lab does not promise
- No production readiness guarantee. A passing playbook is evidence, not approval.
- No compliance or regulatory certification. Verification is not certification.
- No safety validation. Templates and playbooks are planning tools, not safety reviews.
- No model ranking. The lab does not score candidates against each other.
- No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.