Skip to content
WebmasterID

Hub

AI API Pricing

Per-unit API pricing for tracked models. Pricing values are only displayed when sourced from official provider documentation. Cache and batch pricing surface here when published by the vendor.

Cache-pricing semantics differ across providers

We do not collapse provider-specific cache pricing into a single row. Anthropic publishes per-token TTL cache writes (5-minute and 1-hour windows) and per-token cache reads. Google publishes a per-hour cache storage rate alongside a one-shot cache write fee. DeepSeek publishes an input cache-hit rate. The "Cache hit / 1M" column below shows the cache-read (or cache-hit-input) rate where the provider publishes one, and renders Data not yet verified. otherwise.

Hosted-provider pricing is not the same as model-creator pricing

A hosted platform (Groq, Together AI, Bedrock, Vertex, …) may expose a model created by another organisation under a platform-specific model ID, and bill for it at a rate set by the platform — not by the model's creator. /pricing surfaces these as two reference sections: first-party API pricing references (where the billing provider IS the model creator) and hosted provider pricing references (where they differ). Hosted rows are reference values, not a comparison engine — there is no "cheaper" column and no winner. See /research/api-pricing-methodology for the methodology and /docs/pricing-fields for the schema.

7 first-party rows, 2 hosted rows, 5 pending.

First-party API pricing references (7)

Rows where the billing provider is the same organisation that created the model. Volatility on these rows defaults to medium — first-party rates move less than hosted-platform rates but still change (promotional discount windows, regional adjustments, model retirements). Treat each row as a source-backed reference and re-verify against the linked source before projecting cost.

ModelProviderInput / 1MOutput / 1MCache hit / 1MBatch in / 1MPricing contextVolatilityFreshnessSourceLast checked
Claude Opus 4.7Anthropic$5.00$25.00$0.50$2.50First-party APIMediumFreshAnthropic2026-05-20
Claude Sonnet 4.6Anthropic$3.00$15.00$0.30$1.50First-party APIMediumFreshAnthropic2026-05-20
Claude Haiku 4.5Anthropic$1.00$5.00$0.10$0.50First-party APIMediumFreshAnthropic2026-05-20
Gemini 2.5 ProGoogle$1.25$10.00Data not yet verified.$0.63First-party APIMediumFreshGoogle AI2026-05-20
DeepSeek V4 ProDeepSeek$1.74$3.48$0.01Data not yet verified.First-party APIMediumFreshDeepSeek2026-05-21
Mistral Large 3Mistral$0.50$1.50Data not yet verified.Data not yet verified.First-party APIMediumFreshMistral2026-05-21
Claude Opus 4Anthropic$15.00$75.00$1.50$7.50First-party APIMediumFreshAnthropic2026-05-20

Hosted provider pricing references (2)

Rows where a third-party platform hosts a model created by another organisation. The Model creator column is the organisation that built the model; the Billing provider is the platform that invoices for inference. The two are different — and the billing provider sets the rate. Hosted rates carry high volatility by default; re-verify against the platform's own pricing page before any cost projection.

ModelModel creatorBilling providerHosted model IDInput / 1MOutput / 1MCache hit / 1MPricing contextVolatilityFreshnessSourceLast checked
Llama 4 ScoutMetaGroqmeta-llama/llama-4-scout-17b-16e-instruct$0.11$0.34Data not yet verified.Hosted providerHighFreshGroq2026-05-23
DeepSeek V4 ProDeepSeekTogether AIdeepseek-ai/DeepSeek-V4-Pro$2.10$4.40$0.20Hosted providerHighFreshTogether AI2026-05-23

Pending or unavailable creator pricing (5)

Models where no first-party (model-creator) API pricing has been verified. Reasons differ: Meta does not run a paid first-party Llama API at all (hosted pricing may exist on Groq / Together — see above); OpenAI's docs site returns HTTP 403 to automated retrieval; Mistral's pricing tab is JavaScript-driven. Each case is logged on /coverage.

AI API Pricing | WebmasterID Models