Skip to content
WebmasterID

Learn · pricing and hosted

AI model pricing references explained

Why catalogue pricing rows are references, not live quotes — and how to read them without ranking models by price.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

Pricing references are not live quotes

Every pricing row in the catalogue is a reference — a verified field whose value comes from a primary source (the provider's published pricing page) and whose retrieval date is recorded. None of these values are live quotes against the provider's billing system, and none should be treated as a guarantee of what your invoice will say next month.

This matters because AI model pricing changes more often than most infrastructure pricing — sometimes weekly. A pricing row that was correct on its retrieval date can be stale by the time you read it. The catalogue records the date so the reader can judge how confident to be.

First-party vs hosted pricing

The catalogue records two separate kinds of pricing rows:

  • First-party pricing — set by the model creator and published on their official pricing page. This is what you pay when you call the creator's API directly.
  • Hosted pricing — set by a hosting platform for the model it serves. Same model, different provider, different price, different terms. See the lesson on hosted vs first-party.

These rows live in different parts of the data layer. Mixing them in a single ranking is a category error — they come from different invoices.

What to inspect on every pricing row

  • The pricing unit — per million input tokens, per million output tokens, per cached read, per cache write, per second of compute? Units do not compare across providers without alignment.
  • The currency — most rows are USD; some hosting platforms publish in other currencies. Always check.
  • The retrieval date — the catalogue records when the source page was last read. Older retrievals are more likely to be stale.
  • The freshness state — the reverification queue flags rows past the staleness threshold.
  • Whether the row applies to your prompt size — some providers tier pricing on prompt length.
  • Whether the row is first-party or hosted — the catalogue tags this explicitly.

Volatility and freshness

AI model pricing volatility comes from several sources: model launches, snapshot rotations, hosted platform promotions, and currency adjustments. The catalogue records every row with a retrieval date and flags rows that exceed a freshness window in the reverification queue. The queue is the single source of truth for "what needs to be re-read against the provider's page before I trust it."

Why no price ranking

The catalogue deliberately does not publish a least-cost model leaderboard. Pricing semantics differ enough across providers (input vs output tokens, cache reads vs cache writes, per-second compute vs per-request, prompt-size tiers) that a single numeric ranking would mislead more often than it would help. The reader's own traffic pattern decides the total cost — the catalogue gives the per-unit references to model that traffic against, not the ranking it would produce.

Common mistakes

  • Comparing per-token rates without checking units.

    Input vs output, cached vs uncached, and provider-specific accounting can make a numerically smaller rate the more expensive choice for your traffic.

  • Mixing first-party and hosted pricing in one chart.

    They come from different sources and different invoices. Compare within a category, not across.

  • Treating an old retrieval date as 'current'.

    Pricing pages change without notice. The retrievedAt field tells you how old the row is — use it.

  • Reading the absence of a pricing row as 'free'.

    Absence usually means 'no primary-source citation on record' — not that the model is free to use.

Apply this workflow

Apply this workflow

Practise this lesson

These exercises route the lesson concept through the verified-data product surfaces. Each one ends with a concrete artifact you can share.

Data gaps to watch

When a model has known pricing but no row in the catalogue, it usually means the provider does not publish pricing in a primary-source page yet, or the catalogue has not retrieved it. Either case is a verification question, not a missing feature. Open /coverage to see which providers have the thinnest pricing coverage right now.

Related pages

Sources and freshness

Pricing rows carry their own primary-source citations and retrieval dates. Inspect them via the per-model record or via /sources.

Teaching example

Illustrative — not a recommendation.

Situation: Finance asks for a monthly cost projection across three candidate models. The catalogue shows pricing rows with retrievedAt dates that range from two weeks to three months old.

Decision to make: Which pricing rows are fresh enough to use in the projection, and which must be re-verified first?

Verified fields that matter:

  • Pricing reference per candidate (with unit)
  • RetrievedAt date per row
  • Currency
  • Reverification queue entry (if any)
  • First-party vs hosted distinction

What to inspect next:

Weak vs better approach

Weak approach

  • Treat the catalogue's pricing rows as a live invoice quote.
  • Sort candidates by lowest per-token rate.
  • Ignore the retrievedAt date on each row.
  • Mix first-party and hosted rows in the same comparison.

Better approach

  • Treat every row as a sourced reference with a retrievedAt date.
  • Compare on unit semantics, not just numeric values.
  • Re-verify rows older than your freshness threshold before quoting.
  • Keep first-party and hosted rows in separate sections.

Why better: Pricing volatility means even fresh rows go stale quickly. The better approach makes the freshness state part of the projection rather than hiding it.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Pricing-reference note for finance review

## Pricing reference (illustrative)
Candidate: <slug>
Source URL: <provider pricing page>
RetrievedAt: <date>

Per-unit reference (verbatim from the source):
- Input: <amount> per <unit>
- Output: <amount> per <unit>
- Cache (if applicable): <amount> per <unit>

Freshness: <fresh / review-due / stale> (per /reverification)
Caveat: Reference, not a live quote. Re-verify before signing a contract.

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Read pricing rows as references, not invoiceable quotes.

  2. Step 2

    Apply in /pricing →

    Inspect verified pricing rows with retrievedAt dates.

  3. Step 3

    Verify in /reverification →

    Confirm whether any pricing row is flagged stale.

  4. Step 4

    Test in /briefs/build →

    Embed the pricing reference + retrievedAt in the brief.

Review before moving on

  • Every pricing row in my notes has a retrievedAt date.
  • Unit semantics are recorded alongside numeric values.
  • I have NOT ranked candidates by per-token price.
  • Stale rows are listed for /reverification before reuse.
  • First-party and hosted rows are kept in separate sections of the brief.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Ranking models by price — the catalogue records pricing references, not a leaderboard.
  • Ranking which model costs the least for any workload — unit semantics differ, total cost depends on your traffic pattern.
  • Quoting current invoiceable prices — citations age; always re-verify before committing.