Skip to content
WebmasterID

Hub

AI API Pricing

Per-unit API pricing for tracked models. Pricing values are only displayed when sourced from official provider documentation. Cache and batch pricing surface here when published by the vendor.

Cache-pricing semantics differ across providers

We do not collapse provider-specific cache pricing into a single row. Anthropic publishes per-token TTL cache writes (5-minute and 1-hour windows) and per-token cache reads. Google publishes a per-hour cache storage rate alongside a one-shot cache write fee. DeepSeek publishes an input cache-hit rate. The "Cache hit / 1M" column below shows the cache-read (or cache-hit-input) rate where the provider publishes one, and renders Data not yet verified. otherwise.

Hosted-provider pricing is not the same as model-creator pricing

A hosted platform (Groq, Together AI, Bedrock, Vertex, …) may expose a model created by another organisation under a platform-specific model ID, and bill for it at a rate set by the platform — not by the model's creator. /pricing surfaces these as two reference sections: first-party API pricing references (where the billing provider IS the model creator) and hosted provider pricing references (where they differ). Hosted rows are reference values, not a comparison engine — there is no "cheaper" column and no winner. See /research/api-pricing-methodology for the methodology and /docs/pricing-fields for the schema.

Reset0 first-party rows, 1 hosted row, 0 pending.

First-party API pricing references (0)

Rows where the billing provider is the same organisation that created the model. Volatility on these rows defaults to medium — first-party rates move less than hosted-platform rates but still change (promotional discount windows, regional adjustments, model retirements). Treat each row as a source-backed reference and re-verify against the linked source before projecting cost.

No verified pricing rows match the current filters.

Hosted provider pricing references (1)

Rows where a third-party platform hosts a model created by another organisation. The Model creator column is the organisation that built the model; the Billing provider is the platform that invoices for inference. The two are different — and the billing provider sets the rate. Hosted rates carry high volatility by default; re-verify against the platform's own pricing page before any cost projection.

ModelModel creatorBilling providerHosted model IDInput / 1MOutput / 1MCache hit / 1MPricing contextVolatilityFreshnessSourceLast checked
Llama 4 ScoutMetaGroqmeta-llama/llama-4-scout-17b-16e-instruct$0.11$0.34Data not yet verified.Hosted providerHighFreshGroq2026-05-23

Pending or unavailable creator pricing (0)

Models where no first-party (model-creator) API pricing has been verified. Reasons differ: Meta does not run a paid first-party Llama API at all (hosted pricing may exist on Groq / Together — see above); OpenAI's docs site returns HTTP 403 to automated retrieval; Mistral's pricing tab is JavaScript-driven. Each case is logged on /coverage.

No pending rows match the current filters.

AI API Pricing | WebmasterID Models