Skip to content
WebmasterID

Research guide

Output token limits and what they constrain

Why max output tokens is a separate dimension from the context window, how it shapes long-form generation, agentic workflows, and structured output, and which models publish what.

Last updated: 2026-05-21

What max output tokens means

Max output tokens is the maximum number of tokens a single model response can contain. It is independent of the input context window: a 1M-token context model with a 128k-token output limit can read a long document but cannot produce a longer one in a single shot.

The field is recorded separately on each model record with its own verified citation. Where the vendor publishes a synchronous limit and a batch/streaming-beta limit, the catalogue records the synchronous number as the primary value and captures the batch/beta number in the row notes — following the same "durable canonical value" discipline as pricing rows.

Long-form generation

Workloads that produce long documents — full reports, code repositories, structured datasets — are bounded by the output limit, not the context window. A 64k-token output budget is roughly 48,000 words; a 128k-token budget is roughly 96,000 words; the exact text-to-token ratio depends on the tokenizer and the language being generated.

For workloads that exceed the single-call output budget, the two structural options are (a) chunking the generation across multiple calls with a continuation prompt, or (b) switching to the provider's batch API for the higher beta limit. Both have implications for cost (see pricing methodology) and for downstream coherence; neither is captured in the headline number alone.

Agentic loops

Agentic workloads spend the output budget across many tool calls and intermediate reasoning steps. Extended-thinking features available on some models (recorded asextendedThinking / adaptiveThinking on the model record) consume the same output budget as the final answer. A workflow that performs ten tool calls and a long reasoning trace can hit the limit even when the final user-visible response is short.

Structured output

When a model is producing structured output (JSON schema conformance, function-call payloads), token efficiency changes — JSON keys and brackets are individually tokenised and accrue against the output limit. A long array of objects can easily double the apparent length once each property name, string quote, and comma is counted as one or more tokens. The output limit is enforced before schema-conformance retries, so a long structured payload that fails validation can be terminated mid-output and leave the response in a partial-JSON state.

Cost implications

Output tokens are uniformly more expensive than input tokens across the providers in the catalogue. Anthropic Opus 4.7 charges $5 / 1M input vs $25 / 1M output. Gemini 2.5 Pro charges $1.25 / 1M input vs $10 / 1M output (standard tier, ≤200k context). DeepSeek V4 Pro charges $1.74 / 1M cache-miss input vs $3.48 / 1M output. A workload that maximises its output budget per call pays substantially more per request than one that produces short responses.

Batch APIs cut both input and output rates by 50% across the providers tracked. Output-heavy workloads benefit proportionally more from the batch tier than input-heavy ones.

Output-budget pressure by use case

How the output token budget interacts with common workloads
Output pressureFailure modeRecommended budget
Long-form generationFull reports, code repositories, structured datasetsHigh — budget consumed proportionally to deliverable length.Truncation mid-response. Continuation prompts needed; coherence may drift.Match to deliverable; consider batch API for longer-than-64k outputs.
Structured outputJSON schema conformance, function-call payloadsHigh per token of effective payload — keys, brackets, commas all count.Mid-output truncation can leave the JSON unparseable.Reserve 2–3× the visible payload size to cover serialization overhead.
Agentic loopsTool calls + reasoning across many turnsModerate; spread across many short calls rather than one long one.Mid-turn truncation forces re-planning; extended-thinking traces eat the same budget.Keep individual call outputs short; checkpoint state between calls.
Code generationWhole-file or whole-module generationHigh; source code tokenises densely.File truncation produces partial/invalid syntax.Generate in functional units; verify each before continuing.

Verified examples

The four million-token-context models in the catalogue have output limits in a 64k–128k range, with one specific value published per provider. The full list lives under Verified today above, and each entry is linked back to its model page where the citation is visible inline.

What this page assumes is verified

Verified today

Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.

  • Claude Opus 4.7: 128k tokens synchronous; 300k via batch beta

    Verified from the Anthropic Models overview synchronous-Messages-API row, and the Message Batches beta header note.

  • Claude Sonnet 4.6: 64k tokens

    Verified from the Anthropic Models overview.

  • Claude Haiku 4.5: 64k tokens

    Verified from the Anthropic Models overview.

  • Gemini 2.5 Pro: 65,536 tokens

    Verified from ai.google.dev/gemini-api/docs/models/gemini-2.5-pro — listed precisely rather than rounded to 64k.

Honest gaps

Data gaps

Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.

  • DeepSeek V4 Pro max output

    The DeepSeek Models & Pricing page documents context window but does not separately publish max output tokens. The field renders as the canonical unverified-data label on the model page.

  • Mistral max output

    Per-model spec card pages 404 to automated retrieval. Mistral Large 3 max output is unverified.

  • Effective throughput at max output

    Time-to-completion for a 128k-token generation depends on the provider's runtime and the prompt; the catalogue records the headline limit, not throughput.

Continue

Related pages