What a context window actually is
A context window is the maximum number of tokens the model can process in a single request — counting the system prompt, the user message, any retrieved context the application includes, and (depending on the provider's accounting) the generated output. Tokens are not characters and not words; they are sub-word units the model's tokenizer produces.
The catalogue records context windows as verified fields where the value comes directly from the provider's primary documentation. Models with no primary-source citation render the canonical unverified-data label instead of an estimate.
What context window does not guarantee
Context capacity is not capability. A model that accepts a million tokens in a single prompt can still:
- Lose recall on facts buried deep in the input.
- Degrade in instruction-following when the prompt grows past the size used during training.
- Cost dramatically more per request at high context sizes, even when the per-token rate looks flat.
- Hit a separate
max output tokenslimit that is much smaller than the input limit.
These behaviours show up only in workload-specific testing. The catalogue surfaces the verified context window and max output limits; the rest is your evaluation work.
What to inspect before assuming a model fits
- The verified context window — large enough for your prompt + retrieved context + expected output?
- The verified max output tokens — separate field. Always check both.
- Pricing references that specifically apply at large prompt sizes (some providers tier pricing on prompt length).
- The model's lifecycle status — context-window improvements often arrive in newer snapshots; older snapshots may be deprecated.
- Whether the model accepts your input modality (text, image, audio, video) — modality is its own verified field.
Verified examples from the catalogue
Context windows vary widely across the models in the catalogue. The table renders the verified field for several models — inspection only, no ranking.
Verified examples · Context window
Each cell is a verified field with a primary-source citation. Click a model name to read the underlying URL.
| Model | Provider | Context window | Source state |
|---|---|---|---|
| Claude Opus 4.7 | Anthropic | 1,000,000 tokens | verified · citation on record |
| Claude Sonnet 4.6 | Anthropic | 1,000,000 tokens | verified · citation on record |
| Claude Haiku 4.5 | Anthropic | 200,000 tokens | verified · citation on record |
| Gemini 2.5 Pro | 1,048,576 tokens | verified · citation on record | |
| DeepSeek V4 Pro | DeepSeek | 1,000,000 tokens | verified · citation on record |
| Mistral Large 3 | Mistral | 256,000 tokens | verified · citation on record |
Inspection only. The catalogue does not rank these models on the field above.
Relation to cost and latency (without numeric claims)
Large prompts cost more — both in money and in wall-clock latency. Exactly how much more depends on the provider's accounting (input vs output tokens, cache writes vs reads, prompt-size tiers) and on the host (first-party APIs and hosted platforms charge differently). The catalogue does not assert either cost or latency. It surfaces the pricing rows with their citations so you can read the terms yourself, and leaves the latency observation to your own tests.
Common mistakes
Treating max output tokens as the same as context window.
They are separate verified fields. A model can have a 1M-token context window and an 8k-token output cap simultaneously.
Assuming context window growth is monotonic across snapshots.
Newer snapshots usually expand context, but deprecated snapshots may keep older limits. Always check the lifecycle field too.
Counting characters or words instead of tokens.
A token is a sub-word unit. The token count for the same text varies between models because each provider's tokenizer is different.
Apply this workflow
Apply this workflow
Open the selection workspace →
Narrow a source-backed shortlist using verified catalogue fields.
Compare verified fields side by side →
Render up to four models against each other from the typed data layer.
Inspect the citation registry →
Open every primary source the catalogue references for a model or pricing row.
Practise this lesson
These exercises route the lesson concept through the verified-data product surfaces. Each one ends with a concrete artifact you can share.
- beginner8 min
Build your first source-backed shortlist →
Pick a use case, filter the catalogue by verified fields, and end with a shortlist URL you can share with the team.
- beginner7 min
Compare context windows without ranking models →
Use the comparison builder to render verified context window + max output tokens for 3–4 candidate models side by side.
Data gaps to watch
When a model's context window renders as the unverified-data label, the catalogue has not yet recorded a primary-source citation for that value. Confirm externally against the provider's documentation — and consider opening a reverification request from /reverification.
Related pages
- /use-cases/long-context-analysis — the use case that explicitly weights context window.
- /compare/build — render context window alongside max output, pricing, lifecycle.
- /docs/model-page-schema — the data-model definition for context fields.
Sources and freshness
Every context window in the table above is wrapped with a verified field that points at the provider's official model documentation. Citations age over time; the reverification queue shows what is due for re-check.
Teaching example
Illustrative — not a recommendation.
Situation: An application sends 80–200 page PDFs into the model along with a short instruction. The team is trying to decide whether the prompt size is comfortably inside the candidate's verified limits.
Decision to make: Does the prompt + retrieved context + expected output stay inside the candidate's verified context window AND its verified max-output cap?
Verified fields that matter:
- Verified context window
- Verified max output tokens (separate field)
- Pricing reference for the prompt-size tier
- Modality channels (PDF vs image vs text)
What to inspect next:
Weak vs better approach
Weak approach
- Pick the model with the largest published context window.
- Skip max-output verification.
- Assume recall is constant across the window.
- Treat marketing copy as a verified field.
Better approach
- Inspect the verified context window AND the verified max-output cap.
- Confirm the pricing reference for prompt-size tiers, if any.
- Run the long-context testing playbook on representative prompts.
- Record where recall degrades in your evidence brief.
Why better: Context window is a necessary condition, not a sufficient one. The better approach surfaces the workload-specific behaviour the catalogue cannot measure for you.
Example artifact
Illustrative example — not a recommendation. Substitute your own values when you run the workflow.
Long-context shortlist note
## Long-context shortlist (illustrative)
Workload: PDFs averaging 100k tokens + 8k expected output
| Candidate | Context (verified) | Max output (verified) | Source URL | retrievedAt |
| --- | --- | --- | --- | --- |
| <model-A> | <tokens> | <tokens> | <provider docs> | <date> |
| <model-B> | <tokens> | unverified-data | <provider docs> | <date> |
Notes:
- <model-B> has no published max output — flag for external test.
- Both candidates pass the verified context check for our 100k prompts.Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.
Concept → workflow bridge
Step 1
Learn the concept →Understand what context window does and does not guarantee.
Step 2
Apply in /select →Filter the catalogue by verified context window for your workload.
Step 3
Verify in /sources →Read the provider documentation behind each context value.
Step 4
Test in /lab →Run the long-context testing playbook against your candidates.
Review before moving on
- I can name the verified context window for each candidate (or note it is unverified).
- I separately know each candidate's verified max output cap.
- I have not assumed the largest window is automatically the best fit.
- I have a /compare/build URL with the candidates rendered side by side.
- I plan to run the long-context testing playbook before integration.
Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.
What this lesson does not teach
- Predicting how a specific model will behave on a specific prompt — context size is a necessary but not sufficient condition.
- Asserting cost or latency for a long-context prompt — the catalogue does not measure either.
- Picking the model with the largest context window — bigger context does not equal better fit.