Skip to content
WebmasterID

Research guide

Context windows in practice

What a context window actually means for production workloads, why a million-token window on one provider is not the same as a million-token window on another, and how WebmasterID Models verifies the number.

Last updated: 2026-05-21

What a context window actually is

A context window is the maximum number of tokens a single request can carry as input. It is the union of the system prompt, the conversation history, any retrieved documents, tool definitions, and the new user message. A model with a 1,000,000-token context window can accept up to that many tokens in one inference call — not across a session, not across an agent loop, but in one shot.

The catalogue records context window as a verified field on the model entity. Where it is not directly published in the vendor's docs, the record renders the canonical unverified-data label rather than guessing.

Input limit vs output limit

Context window is an input limit. The number of tokens a model can generate in a single response is a separate field — max output tokens — and is usually a much smaller number than the context window. Claude Opus 4.7 advertises a 1M-token context window and a 128k-token max output (300k via the Message Batches beta header). Gemini 2.5 Pro advertises a ~1M-token context and a 65,536-token max output. Mixing these up is a common source of cost and design errors: a workload that needs to read a long document is different from one that needs to write one.

See /research/model-output-limits for the output side.

Tokenizer differences across providers

One million tokens at provider A is not the same payload as one million tokens at provider B. Each provider ships its own tokenizer, and a given fixed text fragments differently depending on which tokenizer encodes it. Anthropic explicitly documents that Opus 4.7 can use up to 35% more tokens than earlier Claude generations for the same input — a fact captured as a verified note on that model's record.

The catalogue does not publish a tokenizer-normalised context window. The headline figures on the model pages are the vendor's own published numbers; a cost projection that needs precision should sample real prompts on each vendor's tokenizer.

Verified context + output matrix

Verified context and output limits across catalogue models
Context (input) tokensMax output tokensSource
Claude Opus 4.7Anthropic Models overview, current flagship1,000,000128,000 (300k via Message Batches beta header)/models/claude-opus-4-7
Claude Sonnet 4.61,000,00064,000/models/claude-sonnet-4-6
Claude Haiku 4.5200,00064,000/models/claude-haiku-4-5
Gemini 2.5 ProPer-model docs at ai.google.dev1,048,57665,536/models/gemini-2-5-pro
DeepSeek V4 ProDeepSeek Models & Pricing1,000,000—/models/deepseek-v4-pro
Claude Opus 4 (deprecated)Retired 2026-06-15200,00032,000/models/claude-opus-4

DeepSeek V4 Pro's max output is not separately published on the Models & Pricing page; the field renders as the canonical unverified-data label on its model page.

Verified examples in the catalogue

The four million-token-class records currently in the catalogue are listed under Verified today above. The contrast between Claude Opus 4 (200k, deprecated) and Claude Opus 4.7 (1M, current) is a useful illustration of how quickly the context-window axis moves within a single provider generation.

The two-sided verified comparison Gemini 2.5 Pro vs Claude Opus 4.7 lays the context-window field side by side along with pricing and modality.

Long-context cost considerations

Long-context requests are sometimes cheaper per token than short ones (with batch APIs), sometimes more expensive (with prompt-size tiers), and sometimes the same. Google's Gemini API charges a surcharge for prompts >200k tokens — captured as the 1M input tokens (>200k context) unit on the pricing hub. Anthropic and DeepSeek currently have a flat per-token rate independent of prompt size.

Provider-specific cache pricing also interacts with long context — see /research/api-pricing-methodology.

What we do not claim

We do not publish a quality-of-recall metric across long contexts (the "needle in a haystack" style of evaluation lives elsewhere and warrants independent replication, not republication). We do not claim a million-token context is equivalent to a million-token context on another provider. We do not publish a latency number for long-context inference — request latency is workload-dependent and we have not yet wired an independent measurement harness.

What this page assumes is verified

Verified today

Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.

  • Anthropic Claude Opus 4.7 and Sonnet 4.6: 1,000,000 token context

    Verified from Anthropic's Models overview. Per the same page, Opus 4.7 ships a new tokenizer that can use up to 35% more tokens for the same text — verified field directly on the model record.

  • Google Gemini 2.5 Pro: 1,048,576 token context

    Verified from ai.google.dev/gemini-api/docs/models/gemini-2.5-pro. Listed precisely as 1,048,576 rather than rounded.

  • DeepSeek V4 Pro: 1,000,000 token context

    Verified from the DeepSeek Models & Pricing page. Same nominal headline figure as the Claude family and Gemini Pro, but different tokenizer and different pricing structure for long prompts.

  • Claude Opus 4 (deprecated): 200,000 token context

    Verified historical record. Demonstrates that within a single provider, generation-to-generation context window changes can be 5×.

Honest gaps

Data gaps

Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.

  • Mistral Large 3 context window

    The per-model spec card on docs.mistral.ai returns 404 to automated retrieval. Mistral context-window claims are out of scope until a manual browser pass lands.

  • Per-prompt effective context

    Providers publish a maximum context length, not an effective one. Effective in-context recall is workload-dependent and is not republished as a metric on this site.

Continue

Related pages