Hub
AI API Pricing
Per-unit API pricing for tracked models. Pricing values are only displayed when sourced from official provider documentation. Cache and batch pricing surface here when published by the vendor.
Cache-pricing semantics differ across providers
We do not collapse provider-specific cache pricing into a single row. Anthropic publishes per-token TTL cache writes (5-minute and 1-hour windows) and per-token cache reads. Google publishes a per-hour cache storage rate alongside a one-shot cache write fee. DeepSeek publishes an input cache-hit rate. The "Cache hit / 1M" column below shows the cache-read (or cache-hit-input) rate where the provider publishes one, and renders Data not yet verified. otherwise.
Hosted-provider pricing is not the same as model-creator pricing
A hosted platform (Groq, Together AI, Bedrock, Vertex, …) may expose a model created by another organisation under a platform-specific model ID, and bill for it at a rate set by the platform — not by the model's creator. /pricing surfaces these as two reference sections: first-party API pricing references (where the billing provider IS the model creator) and hosted provider pricing references (where they differ). Hosted rows are reference values, not a comparison engine — there is no "cheaper" column and no winner. See /research/api-pricing-methodology for the methodology and /docs/pricing-fields for the schema.
First-party API pricing references (7)
Rows where the billing provider is the same organisation that created the model. Volatility on these rows defaults to medium — first-party rates move less than hosted-platform rates but still change (promotional discount windows, regional adjustments, model retirements). Treat each row as a source-backed reference and re-verify against the linked source before projecting cost.
| Model | Provider | Input / 1M | Output / 1M | Cache hit / 1M | Batch in / 1M | Pricing context | Volatility | Freshness | Source | Last checked |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | $0.50 | $2.50 | First-party API | Medium | Fresh | Anthropic | 2026-05-20 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $0.30 | $1.50 | First-party API | Medium | Fresh | Anthropic | 2026-05-20 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | $0.50 | First-party API | Medium | Fresh | Anthropic | 2026-05-20 |
| Gemini 2.5 Pro | $1.25 | $10.00 | Data not yet verified. | $0.63 | First-party API | Medium | Fresh | Google AI | 2026-05-20 | |
| DeepSeek V4 Pro | DeepSeek | $1.74 | $3.48 | $0.01 | Data not yet verified. | First-party API | Medium | Fresh | DeepSeek | 2026-05-21 |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | Data not yet verified. | Data not yet verified. | First-party API | Medium | Fresh | Mistral | 2026-05-21 |
| Claude Opus 4 | Anthropic | $15.00 | $75.00 | $1.50 | $7.50 | First-party API | Medium | Fresh | Anthropic | 2026-05-20 |
Hosted provider pricing references (2)
Rows where a third-party platform hosts a model created by another organisation. The Model creator column is the organisation that built the model; the Billing provider is the platform that invoices for inference. The two are different — and the billing provider sets the rate. Hosted rates carry high volatility by default; re-verify against the platform's own pricing page before any cost projection.
| Model | Model creator | Billing provider | Hosted model ID | Input / 1M | Output / 1M | Cache hit / 1M | Pricing context | Volatility | Freshness | Source | Last checked |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Llama 4 Scout | Meta | Groq | meta-llama/llama-4-scout-17b-16e-instruct | $0.11 | $0.34 | Data not yet verified. | Hosted provider | High | Fresh | Groq | 2026-05-23 |
| DeepSeek V4 Pro | DeepSeek | Together AI | deepseek-ai/DeepSeek-V4-Pro | $2.10 | $4.40 | $0.20 | Hosted provider | High | Fresh | Together AI | 2026-05-23 |
Pending or unavailable creator pricing (5)
Models where no first-party (model-creator) API pricing has been verified. Reasons differ: Meta does not run a paid first-party Llama API at all (hosted pricing may exist on Groq / Together — see above); OpenAI's docs site returns HTTP 403 to automated retrieval; Mistral's pricing tab is JavaScript-driven. Each case is logged on /coverage.
- Verified
- Verified
- DeepSeek R1-0528 (historical)Partial
DeepSeek · provider page
- Mistral Large 2 (retired)Partial
Mistral · provider page
- GPT-5Unverified
OpenAI · provider page