Skip to content
WebmasterID

Reference

Pricing fields reference

Every value the PricingUnit union can take — input, output, cache write (5m / 1h), cache read, per-hour cache storage, batch tiers, prompt-size tiers, the unknown placeholder — with rules for when each may carry a verified amount.

Last updated: 2026-05-21

PricingUnit enum reference

PricingUnit enum reference
UnitFamilyMeaningVerified for
1M input tokensBasePer-million-tokens input rate for synchronous requests.Anthropic, Google, DeepSeek
1M output tokensBasePer-million-tokens output rate for synchronous requests.Anthropic, Google, DeepSeek
1M cache write tokens (5m)CacheAnthropic-style TTL cache write fee — 5-minute window.Anthropic
1M cache write tokens (1h)CacheAnthropic-style TTL cache write fee — 1-hour window.Anthropic
1M cache read tokensCacheAnthropic-style cache read; also used for DeepSeek-style cache-hit input.Anthropic, DeepSeek
1M cache storage / hourCacheGoogle-style per-hour cache storage rate. Independent of cache write fee.NOT interchangeable with the Anthropic TTL units — different semantics.Google
1M input tokens (>200k context)Prompt-size tierSurcharge rate for input on prompts exceeding 200k tokens.Google
1M output tokens (>200k context)Prompt-size tierOutput surcharge for prompts exceeding 200k tokens.Google
1M cache write tokens (>200k context)Prompt-size tierCache write surcharge for prompts exceeding 200k tokens.Google
1M batch input tokensBatchBatch-API input rate (typically 50% of synchronous; higher latency).Anthropic, Google
1M batch output tokensBatchBatch-API output rate (typically 50% of synchronous; higher latency).Anthropic, Google
1M batch input tokens (>200k context)BatchBatch surcharge for inputs exceeding 200k tokens.Google
1M batch output tokens (>200k context)BatchBatch surcharge for outputs exceeding 200k tokens.Google
requestNon-tokenPer-request fee. Rare across the providers tracked.—
imageNon-tokenPer-image fee for image generation or vision-pricing schedules.—
minuteNon-tokenPer-minute fee for audio/transcription products.—
unknownPlaceholderReserved for rows whose unit semantics have not been verified.Integrity guard refuses any row with a verified amount AND this unit.—

The "Verified for" column reflects providers with at least one verified pricing row at that unit today; absence does not mean a provider lacks the concept, only that it has not been verified into the catalogue.

Pricing row shape

interface VerifiedPricingTier {
  unit: PricingUnit;
  amount: MaybeVerified<number>; // verified number in USD, or null
  notes?: string;
}

A row sets exactly one unit from the closed PricingUnit union and a single optional amount. The amount is either a verified number with citation or null; interpolation is never allowed. Each pricing row attaches to one model record under its pricing array.

PricingContext: creator vs host

Sprint 19 added an explicit pricingContext tag on every pricing record so first-party model-creator pricing is never confused with third-party hosted-provider pricing.

type PricingContext =
  | "model_creator_first_party_api"
  | "hosted_provider_api"
  | "cloud_marketplace"
  | "unknown";

interface PricingRecord {
  modelSlug: string;
  modelCreatorProviderSlug: string;   // who made the model
  billingProviderSlug: string;        // who invoices the developer
  hostedModelId?: string;             // platform-specific model ID
  pricingContext: PricingContext;
  tiers: VerifiedPricingTier[];
  citation?: SourceCitation;
  // ...
}
  • model_creator_first_party_api — the billing provider IS the model's creator (Anthropic charging for Claude, DeepSeek charging for DeepSeek V4 Pro, Mistral charging for Mistral Large 3). Existing rows on each ModelEntity.pricing array carry this context implicitly.
  • hosted_provider_api — a third-party platform (Groq, Together AI) hosts a model created by another organisation and bills the developer at the platform's own rate. The two providers are different. hostedModelId records the platform-specific identifier (e.g. Groq's meta-llama/llama-4-scout-17b-16e-instruct).
  • cloud_marketplace — reserved for AWS Bedrock / Vertex / Azure pricing rows. Not used today; reserved so future rows have a stable place to land.
  • unknown — a row whose context has not been determined. Type-level placeholder only; no row currently uses it.

Hosted pricing is intentionally NOT merged into the model creator's schema.org Offer records. Schema.org consumers (search engines, LLMs) treat creator + offers as a single claim — emitting Groq's rate under Meta's creator block would falsely imply Meta charges that rate.

Pricing freshness + volatility

Sprint 20 added two fields that pair with every pricing record to make staleness visible. Pricing values are references, not live quotes — the freshness field tells you how recently a row was confirmed; the volatility field tells you how often the rate is expected to change.

type PricingVolatility = "high" | "medium" | "low" | "unknown";

interface PricingRecord {
  // ...
  volatility: PricingVolatility;
  reviewCadenceDays?: number;
  lastCheckedAt: string | null;
}

// Freshness is computed (lib/pricing-freshness.ts):
type PricingFreshnessState =
  | "fresh"         // checked within 14 days
  | "review_due"    // 15–30 days
  | "stale"         // 31+ days
  | "unknown";      // no timestamp
  • Hosted-provider rows default to volatility "high" and review cadence 14 days — hosting platforms re-price frequently.
  • First-party rows default to "medium" and review cadence 30 days — first-party rates move slower but still move (promotional windows, retirements, regional adjustments).
  • No row ever defaults to "low" volatility. A reader who treats any pricing record as stable should be corrected by every surface that renders it.

Freshness is computed against siteConfig.buildDate rather than wall-clock now so the same build renders the same state on every page. Transitions happen at deploy time, not mid-render. The thresholds (14 / 30 / 45 days) are encoded in PRICING_FRESHNESS_DAYS in lib/pricing-freshness.ts.

Not a live-quote policy: Pricing is a source-backed reference, not a real-time API. Every surface that shows a price also shows the source URL and lastCheckedAt so a reader can audit the row before projecting cost. WebmasterID Models does not rank models or billing providers by price. Rows that age into review_due or staleappear on the /reverification queue with the source URL and a suggested manual action — they are never auto-updated.

Input / output base units

  • "1M input tokens" — per-million-tokens base input rate (USD).
  • "1M output tokens" — per-million-tokens base output rate (USD).

Universal axes; every verified-pricing model has these two.

Cache pricing units

  • "1M cache write tokens (5m)" — Anthropic-style 5-minute TTL cache write.
  • "1M cache write tokens (1h)" — Anthropic-style 1-hour TTL cache write.
  • "1M cache read tokens" — Anthropic-style cache read; also used for DeepSeek-style cache-hit input where the provider publishes a per-token rate.
  • "1M cache storage / hour" — Google-style per-hour cache storage rate. NOT interchangeable with the Anthropic TTL units; recorded distinctly so cost projections do not conflate them.

Prompt-size tier units

  • "1M input tokens (>200k context)"
  • "1M output tokens (>200k context)"
  • "1M cache write tokens (>200k context)"

Used by Google's Gemini API for the surcharge that applies to prompts >200k tokens. Surfaced as first-class units so long-context cost projections can filter on the right number without parsing row notes.

Batch pricing units

  • "1M batch input tokens"
  • "1M batch output tokens"
  • "1M batch input tokens (>200k context)"
  • "1M batch output tokens (>200k context)"

Batch API tiers across providers. Typically 50% of the synchronous rate at every provider tracked, but the catalogue records the published amount, not the derived ratio.

Non-token units

  • "request" — per-request rate (rare).
  • "image" — per-image rate, for image generation or vision-pricing schedules that are not token-denominated.
  • "minute" — per-minute rate for audio/transcription products.

The unknown placeholder

"unknown" is a placeholder for rows whose unit semantics have not yet been verified. A row with this unit MUST NOT carry a verified amount — an integrity guard refuses to ship a build that violates this invariant. The placeholder lets a pricing intent be tracked structurally without distorting a provider's pricing model into another provider's shape.

Validation rules

  • Every pricing row with a verified amount carries a citation.
  • A pricing row with unit:"unknown" may not carry a verified amount.
  • DeepSeek first-party pricing must be anchored to a DeepSeek citation. (Same per-provider citation discipline applies elsewhere.)
  • Every hosted-provider pricing record has a different modelCreatorProviderSlug and billingProviderSlug; first-party records keep them equal.
  • Hosted-provider rows must cite the billing provider's own pricing page — Groq prices must be sourced from Groq, Together prices from Together. A Meta citation may never sit on a Groq row.
  • Meta first-party Llama pricing remains empty unless an official Meta pricing citation appears in data/citations.ts. Groq and Together pricing for Llama models lives in data/hosted-pricing.ts and does not back-fill Meta's creator pricing.
  • Promotional discount windows are recorded in row notes; the durable canonical amount is the regular rate, not the discounted rate.

See /research/api-pricing-methodology for the reasoning behind these rules.

Continue

Related pages