Research guide
Output token limits and what they constrain
Why max output tokens is a separate dimension from the context window, how it shapes long-form generation, agentic workflows, and structured output, and which models publish what.
Last updated: 2026-05-21
What max output tokens means
Max output tokens is the maximum number of tokens a single model response can contain. It is independent of the input context window: a 1M-token context model with a 128k-token output limit can read a long document but cannot produce a longer one in a single shot.
The field is recorded separately on each model record with its own verified citation. Where the vendor publishes a synchronous limit and a batch/streaming-beta limit, the catalogue records the synchronous number as the primary value and captures the batch/beta number in the row notes — following the same "durable canonical value" discipline as pricing rows.
Long-form generation
Workloads that produce long documents — full reports, code repositories, structured datasets — are bounded by the output limit, not the context window. A 64k-token output budget is roughly 48,000 words; a 128k-token budget is roughly 96,000 words; the exact text-to-token ratio depends on the tokenizer and the language being generated.
For workloads that exceed the single-call output budget, the two structural options are (a) chunking the generation across multiple calls with a continuation prompt, or (b) switching to the provider's batch API for the higher beta limit. Both have implications for cost (see pricing methodology) and for downstream coherence; neither is captured in the headline number alone.
Agentic loops
Agentic workloads spend the output budget across many tool calls and intermediate reasoning steps. Extended-thinking features available on some models (recorded asextendedThinking / adaptiveThinking on the model record) consume the same output budget as the final answer. A workflow that performs ten tool calls and a long reasoning trace can hit the limit even when the final user-visible response is short.
Structured output
When a model is producing structured output (JSON schema conformance, function-call payloads), token efficiency changes — JSON keys and brackets are individually tokenised and accrue against the output limit. A long array of objects can easily double the apparent length once each property name, string quote, and comma is counted as one or more tokens. The output limit is enforced before schema-conformance retries, so a long structured payload that fails validation can be terminated mid-output and leave the response in a partial-JSON state.
Cost implications
Output tokens are uniformly more expensive than input tokens across the providers in the catalogue. Anthropic Opus 4.7 charges $5 / 1M input vs $25 / 1M output. Gemini 2.5 Pro charges $1.25 / 1M input vs $10 / 1M output (standard tier, ≤200k context). DeepSeek V4 Pro charges $1.74 / 1M cache-miss input vs $3.48 / 1M output. A workload that maximises its output budget per call pays substantially more per request than one that produces short responses.
Batch APIs cut both input and output rates by 50% across the providers tracked. Output-heavy workloads benefit proportionally more from the batch tier than input-heavy ones.
Output-budget pressure by use case
| Output pressure | Failure mode | Recommended budget | |
|---|---|---|---|
| Long-form generationFull reports, code repositories, structured datasets | High — budget consumed proportionally to deliverable length. | Truncation mid-response. Continuation prompts needed; coherence may drift. | Match to deliverable; consider batch API for longer-than-64k outputs. |
| Structured outputJSON schema conformance, function-call payloads | High per token of effective payload — keys, brackets, commas all count. | Mid-output truncation can leave the JSON unparseable. | Reserve 2–3× the visible payload size to cover serialization overhead. |
| Agentic loopsTool calls + reasoning across many turns | Moderate; spread across many short calls rather than one long one. | Mid-turn truncation forces re-planning; extended-thinking traces eat the same budget. | Keep individual call outputs short; checkpoint state between calls. |
| Code generationWhole-file or whole-module generation | High; source code tokenises densely. | File truncation produces partial/invalid syntax. | Generate in functional units; verify each before continuing. |
Verified examples
The four million-token-context models in the catalogue have output limits in a 64k–128k range, with one specific value published per provider. The full list lives under Verified today above, and each entry is linked back to its model page where the citation is visible inline.
What this page assumes is verified
Verified today
Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.
Claude Opus 4.7: 128k tokens synchronous; 300k via batch beta
Verified from the Anthropic Models overview synchronous-Messages-API row, and the Message Batches beta header note.
Claude Sonnet 4.6: 64k tokens
Verified from the Anthropic Models overview.
Claude Haiku 4.5: 64k tokens
Verified from the Anthropic Models overview.
Gemini 2.5 Pro: 65,536 tokens
Verified from ai.google.dev/gemini-api/docs/models/gemini-2.5-pro — listed precisely rather than rounded to 64k.
Honest gaps
Data gaps
Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.
DeepSeek V4 Pro max output
The DeepSeek Models & Pricing page documents context window but does not separately publish max output tokens. The field renders as the canonical unverified-data label on the model page.
Mistral max output
Per-model spec card pages 404 to automated retrieval. Mistral Large 3 max output is unverified.
Effective throughput at max output
Time-to-completion for a 128k-token generation depends on the provider's runtime and the prompt; the catalogue records the headline limit, not throughput.
Continue