Use case
Hosted inference
When the model creator does not run a paid first-party API and inference is delivered by a third-party hosting platform (Groq, Together AI). Hosted availability is a stable identity claim; hosted pricing is a volatile reference value.
Verified fields used
- hosted availability record (host × model)
- hosted model ID
- billing provider slug
- model creator slug
- hosted pricing references + freshness
Hosted inference is the right use case when the model creator does not run a paid first-party API. Meta is the canonical example today: Llama 4 Scout and Llama 4 Maverick are open-weights releases with no Meta first-party API; inference is delivered by third-party hosting platforms (Groq, Together AI) at rates each platform sets independently.
The catalogue separates two distinct claims here:
- Hosted availability is a stable identity fact — the platform exposes the model under a specific hosted model ID (e.g. Groq's
meta-llama/llama-4-scout-17b-16e-instruct). Availability rarely changes and is recorded inlib/hosted-availability.ts. - Hosted pricing is a volatile reference value — the platform's published per-token rate on a specific date. Hosted pricing carries a freshness chip and a high-volatility tag because hosting platforms re-price frequently. See /research/api-pricing-methodology for the full distinction.
Two read-traps to avoid:
- A hosting platform is not the model creator. Groq does not become "the maker of Llama 4 Scout" by exposing it. The catalogue refuses to attribute creator status to a hosting platform; the integrity guards block any such drift.
- Hosted pricing is not creator pricing. Groq's Llama 4 Scout rate reflects Groq's pricing decision — not Meta's. WebmasterID Models does not rank hosting platforms by price, and you should not infer that Together's rate is "cheaper" or "more expensive" in a meaningful sense without a workload-specific projection.
The shortlist below collects models with verified hosted availability records. Each row links to the hosting platform's provider page, where the availability sidebar + hosted pricing references render side-by-side with the verification + freshness state. Re-verify against the platform's own pricing page before any procurement decision.
Shortlist
Shortlist for hosted inference (2)
Generated from the typed local data layer. Shortlist order: verified field count → active lifecycle → source count → name. Not a ranking.
Open comparison builder for this use case → Pre-seeded with the top 4 shortlist candidates; these are candidates, not picks.
Create evidence brief for this use case → Markdown or JSON export of the same selection — evidence, not a recommendation.
| Model | Provider | Lifecycle | Context | Modalities | Sources | Freshness | Next action |
|---|---|---|---|---|---|---|---|
Llama 4 ScoutVerified | meta | active | 10,000,000 | text-in, image-in, text-out | 1 | Fresh | Inspect the hosted-availability record and the third-party hosted pricing reference on the provider page. |
DeepSeek V4 ProVerified | deepseek | active | 1,000,000 | — | 3 | Fresh | Review the verified pricing reference and re-check the freshness state before projecting cost. |