Skip to content
WebmasterID

Research guide

API pricing methodology: what the rows on /pricing mean

How WebmasterID Models tracks AI API pricing — input tokens, output tokens, cache write windows, cache reads, per-hour cache storage, batch tiers — and why provider pricing cannot always be normalised into a single number.

Last updated: 2026-05-21

What a row on /pricing means

A row on the pricing hub is a single tuple — model, unit, amount, citation, last-checked — recorded straight from the vendor's pricing page. Every row is one of the values in the typed PricingUnit union (see /docs/pricing-fields for the full enum). If the row carries a value, it carries a citation; if it carries the canonical unverified-data label, it carries no citation by design.

References, not live quotes

Pricing is volatile. Vendors run promotional discount windows (DeepSeek's 75% v4-pro window), hosting platforms re-price weekly, marketplace rates drift with capacity. Without a freshness signal a price catalogue degrades silently as rows age — readers see numbers and assume they are current. Sprint 20 reframes every pricing surface as a source-backed reference: a value confirmed against a vendor page on a specific date, with an explicit volatility tag and an explicit freshness state.

A row that shows $0.11/$0.34 today is not a guarantee the same rate is in effect tomorrow. Every row links back to the vendor's own page; cost projections should re-verify against the source before commitment. The schema definition lives at /docs/pricing-fields; the rendering layer pairs every row with Freshness and Volatility chips so the warning is unavoidable. Sprint 21 added the manual review loop: rows that age into review_due or stale show up on the /reverification queue with the source URL and a suggested manual action. Nothing on the catalogue auto-fetches or auto-mutates a pricing value.

No price-ranking policy

WebmasterID Models does not rank models or billing providers by price. Pricing is rendered side-by-side where the underlying data is verified, but the surfaces deliberately omit:

  • "Cheapest provider" rankings.
  • "Lower / lowest" comparators or "savings" calculations.
  • "Better value" or "best price" copy.
  • Price-derived winner declarations on comparison pages.
  • Delta columns that compute one row minus another.

The reasons are practical. (1) Tokenizers differ between providers; the same string costs different numbers of tokens on each vendor — a $/1M comparison is meaningful only at the level of a real workload distribution. (2) Cache pricing semantics are not commensurable across vendors (see the cache section below). (3) Promotional windows expire silently — a "cheapest" ranking pinned today is wrong tomorrow. (4) Hosted pricing depends on platform, region, tier, and contract; a public-listing $/1M number is one slice of a multi-variable surface. The data is honest about its limitations rather than papering over them with a derived number that invites bad decisions.

An integrity guard enforces this at build time — any new copy containing "cheapest", "lower cost", "lowest price", "best value", "price winner", "save money", or "cheaper than" fails CI on data + content surfaces.

Model creator vs hosted provider

A pricing row is two-sided: an organisation creates a model, and an organisation bills for inference. The two are usually the same — Anthropic, Google, DeepSeek, and Mistral all run first-party APIs for the models they create — but they are not always the same. Meta releases Llama 4 as open weights and runs no paid API for it; the cost of Llama 4 inference comes from third-party hosting platforms (Groq, Together AI, Bedrock, Vertex, …) at rates each platform sets independently.

Sprint 19 (2026-05-23) made this split explicit. Every pricing record now carries a pricingContext tag:

  • model_creator_first_party_api — the billing provider IS the model's creator. Every existing Anthropic, Google, DeepSeek, and Mistral row is classified here.
  • hosted_provider_api — a hosting platform bills for a model created by someone else. Groq's Llama 4 Scout row and Together AI's DeepSeek V4 Pro row are the two such rows recorded today.

The /pricing hub renders these two contexts in separate tables. Hosted rows include both the Model creator and Billing provider columns so a reader can never mistake a Groq rate for a Meta rate. Hosted pricing also has its own hosted model ID column — the platform-specific identifier the developer actually passes in the API request.

Why this distinction matters for cost projection:

  • Hosted-provider rates do not reflect a model creator's pricing decision; they reflect the hosting platform's. A comparison that uses Groq's Llama 4 Scout rate as "Meta's price" is wrong.
  • Different hosting platforms price the same model differently. Llama 4 Scout on Groq is not Llama 4 Scout on Together; DeepSeek V4 Pro on Together is not DeepSeek V4 Pro on DeepSeek's own API. The two-table layout makes the comparison side-by-side, not collapsed.
  • Promotional discounts and platform-specific features (Groq's prompt-cache discount, Together's cache-hit input rate) are attached to the hosting platform's row, not back-fitted onto the creator's row.

The full schema lives in /docs/pricing-fields.

Hosted availability vs hosted pricing

Sprint 20 separated availability from pricing. Availability is a stable fact — the platform either exposes the model or it does not. Pricing is the volatile reference value — it may change after retrieval and is paired with a freshness signal. The two surface independently in the catalogue so the stable identity claim ("Llama 4 Scout is hosted on Groq under model ID meta-llama/llama-4-scout-17b-16e-instruct") does not degrade in lockstep with the volatile rate.

The availability layer lives in lib/hosted-availability.ts. It exposes one record per (host × model) carrying the hosted model ID, model creator slug, billing provider slug, the underlying citation, the last-checked timestamp, and a computed pricing-freshness state. Provider pages render the availability list above the pricing table; model pages render it alongside hosted pricing references. The catalogue intentionally does not promote availability into a dedicated indexable route unless and until the data is wide enough to warrant one — see the route-decision note in this sprint's commit.

Input and output tokens

The two universal axes. 1M input tokens is the per-million-token price the provider charges for the prompt; 1M output tokens is the same for generated content. These two units are stable across providers and are usually the first comparison axis for cost projection.

Token-to-token equivalence across providers is not exact — a provider's tokenizer determines how many tokens a given string costs. Anthropic Opus 4.7 in particular ships a new tokenizer that can use up to ~35% more tokens for the same fixed text than earlier Claude generations; the verified field on that model's record captures this explicitly. The pricing hub does not adjust for tokenizer differences; cost projections that need precision should sample real prompts on each vendor's tokenizer.

Pricing unit matrix

Every row on /pricing is one of the units below. The matrix documents the vocabulary; per-provider verified amounts live on each model record and render through the existing pricing helpers.

PricingUnit enum reference
UnitFamilyMeaningVerified for
1M input tokensBasePer-million-tokens input rate for synchronous requests.Anthropic, Google, DeepSeek
1M output tokensBasePer-million-tokens output rate for synchronous requests.Anthropic, Google, DeepSeek
1M cache write tokens (5m)CacheAnthropic-style TTL cache write fee — 5-minute window.Anthropic
1M cache write tokens (1h)CacheAnthropic-style TTL cache write fee — 1-hour window.Anthropic
1M cache read tokensCacheAnthropic-style cache read; also used for DeepSeek-style cache-hit input.Anthropic, DeepSeek
1M cache storage / hourCacheGoogle-style per-hour cache storage rate. Independent of cache write fee.NOT interchangeable with the Anthropic TTL units — different semantics.Google
1M input tokens (>200k context)Prompt-size tierSurcharge rate for input on prompts exceeding 200k tokens.Google
1M output tokens (>200k context)Prompt-size tierOutput surcharge for prompts exceeding 200k tokens.Google
1M cache write tokens (>200k context)Prompt-size tierCache write surcharge for prompts exceeding 200k tokens.Google
1M batch input tokensBatchBatch-API input rate (typically 50% of synchronous; higher latency).Anthropic, Google
1M batch output tokensBatchBatch-API output rate (typically 50% of synchronous; higher latency).Anthropic, Google
1M batch input tokens (>200k context)BatchBatch surcharge for inputs exceeding 200k tokens.Google
1M batch output tokens (>200k context)BatchBatch surcharge for outputs exceeding 200k tokens.Google
requestNon-tokenPer-request fee. Rare across the providers tracked.—
imageNon-tokenPer-image fee for image generation or vision-pricing schedules.—
minuteNon-tokenPer-minute fee for audio/transcription products.—
unknownPlaceholderReserved for rows whose unit semantics have not been verified.Integrity guard refuses any row with a verified amount AND this unit.—

Provider cache semantics side-by-side

Cache pricing semantics across providers
AnthropicGoogle GeminiDeepSeek
Cache write feeCost when a payload first enters the cacheTwo TTL tiers: 5-minute and 1-hour cache write rates, both per-million-tokens.Single one-shot cache write fee, per-million-tokens. Independent of storage rate.No explicit write fee published — input rate is split into cache-miss vs cache-hit instead.
Cache read / hitCost when a request reuses cached contentSingle per-million-tokens cache read rate. Same number across all TTLs.No per-read fee — cached tokens are billed via the per-hour storage rate.Cache-hit input rate is dramatically lower than cache-miss input (sometimes ~100× lower).
Storage / retentionRecurring cost while cache existsTTL-bound: cache disappears after 5 minutes or 1 hour depending on which write rate was used.Per-hour storage rate, per-million-tokens, continuing as long as the cache exists.Not disclosed — DeepSeek treats cache state as an implementation detail.

These semantics are deliberately kept as separate units in the catalogue. A cost projection that maps Google's per-hour storage onto Anthropic's TTL writes will be wrong by a multiplier; the only safe way to compare is per-workload arithmetic.

Cache pricing is provider-specific

Cache pricing is where providers diverge most. We deliberately do not collapse the divergent semantics into a single column.

  • Anthropic. Two cache write rates per model (5-minute and 1-hour TTL) plus a cache read rate, all per-million-tokens. The 5-minute write is cheaper; the 1-hour write is more expensive but persists across long sessions.
  • Google. A one-shot cache write fee per million tokens plus a separate per-hour cache storage rate. Storage is a recurring cost while the cache exists; write is a one-time fee per cached payload.
  • DeepSeek. Two input rates — cache-miss and cache-hit — published as separate amounts. The cache-hit rate is dramatically lower than cache-miss; the catalogue records both.

The PricingUnit enum captures these distinctions as separate rows: 1M cache write tokens (5m), 1M cache write tokens (1h), 1M cache read tokens, 1M cache storage / hour. A cost model that ignores the semantic difference will mis-estimate by a wide margin on cache-heavy workloads.

Batch API pricing

Batch API tiers are the cheapest token rate at every provider we track — typically 50% of the synchronous rate. Recorded as 1M batch input tokens and 1M batch output tokens. Latency is much higher and the queue semantics are provider-specific; the pricing hub records the rate, not the SLA.

Prompt-size tiers

Google publishes a two-tier prompt-size price for Gemini: a standard rate for prompts ≤200k tokens and a surcharge for prompts >200k tokens. Both the standard and surcharge are recorded as separate first-class rows (1M input tokens (>200k context), 1M output tokens (>200k context), etc.) so a long-context cost projection can use the right number directly. We did not hide the surcharge in a row note; surfacing it as a unit makes it filterable and citable.

Why we do not normalise to a single "total cost" column

A single "effective $/1M tokens" metric requires assumptions about cache-hit rate, batch share, prompt size distribution, and tokenizer differences — none of which are intrinsic to the model. Publishing a normalised number invites spurious precision and is the exact pattern we are trying to avoid. Readers who need a normalised projection should compute it against their own workload distribution; the pricing hub provides each input to that computation as a separate, sourced row.

Citation rules

Every pricing row with a verified amount carries a primary-source citation, surfaced both inline on the model page and in the source index at /sources. An integrity guard refuses to ship a build where a verified pricing row references no citation. A second guard refuses to ship a row that carries unit unknown alongside a verified amount — the placeholder is for rows whose unit semantics have not yet been confirmed, and such a row may never carry a value.

Why some pricing is hidden

Where a pricing page renders only with JavaScript (as is the case with mistral.ai/pricing's API tab) or returns HTTP 403 to automated retrieval (as platform.openai.com does), no rate is recorded. The relevant model pages render the canonical unverified-data label and the pricing hub lists those models under Pending verification. A manual browser pass is how a value moves out of that bucket; the workflow is documented at /docs/data-verification.

What this page assumes is verified

Verified today

Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.

  • Anthropic pricing rows

    Base input/output, 5-minute and 1-hour cache write, cache read, and batch input/output for every active Claude model are verified from the Anthropic Pricing reference.

  • Google Gemini pricing rows

    Standard tier (≤200k context), prompt-size surcharge (>200k context), batch tier, per-hour cache storage, and one-shot cache write are verified from ai.google.dev/pricing for Gemini 2.5 Pro.

  • DeepSeek V4 Pro pricing rows

    Cache-miss input, cache-hit input, and output rates verified from the DeepSeek Models & Pricing page. The 75% promotional discount window is recorded per row's notes rather than as the durable canonical value.

Honest gaps

Data gaps

Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.

  • Mistral pricing

    mistral.ai/pricing renders Le Chat subscription plans by default; the per-model API pricing tab is JS-driven. Mistral pricing rows are deferred to a manual browser pass.

  • OpenAI pricing

    platform.openai.com returns 403 to automated retrieval. No OpenAI pricing has been verified.

  • Provider promotional rates beyond the documented discount window

    When a provider runs a time-limited discount, the catalogue records the regular rate as the durable canonical value and captures the effective discounted rate in the row's notes. We do not pin a soon-to-expire price as the headline number.

Continue

Related pages