Reference
Pricing fields reference
Every value the PricingUnit union can take — input, output, cache write (5m / 1h), cache read, per-hour cache storage, batch tiers, prompt-size tiers, the unknown placeholder — with rules for when each may carry a verified amount.
Last updated: 2026-05-21
PricingUnit enum reference
| Unit | Family | Meaning | Verified for |
|---|---|---|---|
| 1M input tokens | Base | Per-million-tokens input rate for synchronous requests. | Anthropic, Google, DeepSeek |
| 1M output tokens | Base | Per-million-tokens output rate for synchronous requests. | Anthropic, Google, DeepSeek |
| 1M cache write tokens (5m) | Cache | Anthropic-style TTL cache write fee — 5-minute window. | Anthropic |
| 1M cache write tokens (1h) | Cache | Anthropic-style TTL cache write fee — 1-hour window. | Anthropic |
| 1M cache read tokens | Cache | Anthropic-style cache read; also used for DeepSeek-style cache-hit input. | Anthropic, DeepSeek |
| 1M cache storage / hour | Cache | Google-style per-hour cache storage rate. Independent of cache write fee.NOT interchangeable with the Anthropic TTL units — different semantics. | |
| 1M input tokens (>200k context) | Prompt-size tier | Surcharge rate for input on prompts exceeding 200k tokens. | |
| 1M output tokens (>200k context) | Prompt-size tier | Output surcharge for prompts exceeding 200k tokens. | |
| 1M cache write tokens (>200k context) | Prompt-size tier | Cache write surcharge for prompts exceeding 200k tokens. | |
| 1M batch input tokens | Batch | Batch-API input rate (typically 50% of synchronous; higher latency). | Anthropic, Google |
| 1M batch output tokens | Batch | Batch-API output rate (typically 50% of synchronous; higher latency). | Anthropic, Google |
| 1M batch input tokens (>200k context) | Batch | Batch surcharge for inputs exceeding 200k tokens. | |
| 1M batch output tokens (>200k context) | Batch | Batch surcharge for outputs exceeding 200k tokens. | |
| request | Non-token | Per-request fee. Rare across the providers tracked. | — |
| image | Non-token | Per-image fee for image generation or vision-pricing schedules. | — |
| minute | Non-token | Per-minute fee for audio/transcription products. | — |
| unknown | Placeholder | Reserved for rows whose unit semantics have not been verified.Integrity guard refuses any row with a verified amount AND this unit. | — |
The "Verified for" column reflects providers with at least one verified pricing row at that unit today; absence does not mean a provider lacks the concept, only that it has not been verified into the catalogue.
Pricing row shape
interface VerifiedPricingTier {
unit: PricingUnit;
amount: MaybeVerified<number>; // verified number in USD, or null
notes?: string;
}A row sets exactly one unit from the closed PricingUnit union and a single optional amount. The amount is either a verified number with citation or null; interpolation is never allowed. Each pricing row attaches to one model record under its pricing array.
PricingContext: creator vs host
Sprint 19 added an explicit pricingContext tag on every pricing record so first-party model-creator pricing is never confused with third-party hosted-provider pricing.
type PricingContext =
| "model_creator_first_party_api"
| "hosted_provider_api"
| "cloud_marketplace"
| "unknown";
interface PricingRecord {
modelSlug: string;
modelCreatorProviderSlug: string; // who made the model
billingProviderSlug: string; // who invoices the developer
hostedModelId?: string; // platform-specific model ID
pricingContext: PricingContext;
tiers: VerifiedPricingTier[];
citation?: SourceCitation;
// ...
}model_creator_first_party_api— the billing provider IS the model's creator (Anthropic charging for Claude, DeepSeek charging for DeepSeek V4 Pro, Mistral charging for Mistral Large 3). Existing rows on eachModelEntity.pricingarray carry this context implicitly.hosted_provider_api— a third-party platform (Groq, Together AI) hosts a model created by another organisation and bills the developer at the platform's own rate. The two providers are different.hostedModelIdrecords the platform-specific identifier (e.g. Groq'smeta-llama/llama-4-scout-17b-16e-instruct).cloud_marketplace— reserved for AWS Bedrock / Vertex / Azure pricing rows. Not used today; reserved so future rows have a stable place to land.unknown— a row whose context has not been determined. Type-level placeholder only; no row currently uses it.
Hosted pricing is intentionally NOT merged into the model creator's schema.org Offer records. Schema.org consumers (search engines, LLMs) treat creator + offers as a single claim — emitting Groq's rate under Meta's creator block would falsely imply Meta charges that rate.
Pricing freshness + volatility
Sprint 20 added two fields that pair with every pricing record to make staleness visible. Pricing values are references, not live quotes — the freshness field tells you how recently a row was confirmed; the volatility field tells you how often the rate is expected to change.
type PricingVolatility = "high" | "medium" | "low" | "unknown";
interface PricingRecord {
// ...
volatility: PricingVolatility;
reviewCadenceDays?: number;
lastCheckedAt: string | null;
}
// Freshness is computed (lib/pricing-freshness.ts):
type PricingFreshnessState =
| "fresh" // checked within 14 days
| "review_due" // 15–30 days
| "stale" // 31+ days
| "unknown"; // no timestamp
- Hosted-provider rows default to volatility
"high"and review cadence14days — hosting platforms re-price frequently. - First-party rows default to
"medium"and review cadence30days — first-party rates move slower but still move (promotional windows, retirements, regional adjustments). - No row ever defaults to
"low"volatility. A reader who treats any pricing record as stable should be corrected by every surface that renders it.
Freshness is computed against siteConfig.buildDate rather than wall-clock now so the same build renders the same state on every page. Transitions happen at deploy time, not mid-render. The thresholds (14 / 30 / 45 days) are encoded in PRICING_FRESHNESS_DAYS in lib/pricing-freshness.ts.
Not a live-quote policy: Pricing is a source-backed reference, not a real-time API. Every surface that shows a price also shows the source URL and lastCheckedAt so a reader can audit the row before projecting cost. WebmasterID Models does not rank models or billing providers by price. Rows that age into review_due or staleappear on the /reverification queue with the source URL and a suggested manual action — they are never auto-updated.
Input / output base units
"1M input tokens"— per-million-tokens base input rate (USD)."1M output tokens"— per-million-tokens base output rate (USD).
Universal axes; every verified-pricing model has these two.
Cache pricing units
"1M cache write tokens (5m)"— Anthropic-style 5-minute TTL cache write."1M cache write tokens (1h)"— Anthropic-style 1-hour TTL cache write."1M cache read tokens"— Anthropic-style cache read; also used for DeepSeek-style cache-hit input where the provider publishes a per-token rate."1M cache storage / hour"— Google-style per-hour cache storage rate. NOT interchangeable with the Anthropic TTL units; recorded distinctly so cost projections do not conflate them.
Prompt-size tier units
"1M input tokens (>200k context)""1M output tokens (>200k context)""1M cache write tokens (>200k context)"
Used by Google's Gemini API for the surcharge that applies to prompts >200k tokens. Surfaced as first-class units so long-context cost projections can filter on the right number without parsing row notes.
Batch pricing units
"1M batch input tokens""1M batch output tokens""1M batch input tokens (>200k context)""1M batch output tokens (>200k context)"
Batch API tiers across providers. Typically 50% of the synchronous rate at every provider tracked, but the catalogue records the published amount, not the derived ratio.
Non-token units
"request"— per-request rate (rare)."image"— per-image rate, for image generation or vision-pricing schedules that are not token-denominated."minute"— per-minute rate for audio/transcription products.
The unknown placeholder
"unknown" is a placeholder for rows whose unit semantics have not yet been verified. A row with this unit MUST NOT carry a verified amount — an integrity guard refuses to ship a build that violates this invariant. The placeholder lets a pricing intent be tracked structurally without distorting a provider's pricing model into another provider's shape.
Validation rules
- Every pricing row with a verified amount carries a citation.
- A pricing row with
unit:"unknown"may not carry a verified amount. - DeepSeek first-party pricing must be anchored to a DeepSeek citation. (Same per-provider citation discipline applies elsewhere.)
- Every hosted-provider pricing record has a different
modelCreatorProviderSlugandbillingProviderSlug; first-party records keep them equal. - Hosted-provider rows must cite the billing provider's own pricing page — Groq prices must be sourced from Groq, Together prices from Together. A Meta citation may never sit on a Groq row.
- Meta first-party Llama pricing remains empty unless an official Meta pricing citation appears in
data/citations.ts. Groq and Together pricing for Llama models lives indata/hosted-pricing.tsand does not back-fill Meta's creator pricing. - Promotional discount windows are recorded in row
notes; the durable canonical amount is the regular rate, not the discounted rate.
See /research/api-pricing-methodology for the reasoning behind these rules.
Continue