Skip to content
WebmasterID

Use case

Hosted inference

When the model creator does not run a paid first-party API and inference is delivered by a third-party hosting platform (Groq, Together AI). Hosted availability is a stable identity claim; hosted pricing is a volatile reference value.

Verified fields used

  • hosted availability record (host × model)
  • hosted model ID
  • billing provider slug
  • model creator slug
  • hosted pricing references + freshness

Hosted inference is the right use case when the model creator does not run a paid first-party API. Meta is the canonical example today: Llama 4 Scout and Llama 4 Maverick are open-weights releases with no Meta first-party API; inference is delivered by third-party hosting platforms (Groq, Together AI) at rates each platform sets independently.

The catalogue separates two distinct claims here:

  • Hosted availability is a stable identity fact — the platform exposes the model under a specific hosted model ID (e.g. Groq's meta-llama/llama-4-scout-17b-16e-instruct). Availability rarely changes and is recorded in lib/hosted-availability.ts.
  • Hosted pricing is a volatile reference value — the platform's published per-token rate on a specific date. Hosted pricing carries a freshness chip and a high-volatility tag because hosting platforms re-price frequently. See /research/api-pricing-methodology for the full distinction.

Two read-traps to avoid:

  • A hosting platform is not the model creator. Groq does not become "the maker of Llama 4 Scout" by exposing it. The catalogue refuses to attribute creator status to a hosting platform; the integrity guards block any such drift.
  • Hosted pricing is not creator pricing. Groq's Llama 4 Scout rate reflects Groq's pricing decision — not Meta's. WebmasterID Models does not rank hosting platforms by price, and you should not infer that Together's rate is "cheaper" or "more expensive" in a meaningful sense without a workload-specific projection.

The shortlist below collects models with verified hosted availability records. Each row links to the hosting platform's provider page, where the availability sidebar + hosted pricing references render side-by-side with the verification + freshness state. Re-verify against the platform's own pricing page before any procurement decision.

Shortlist

Shortlist for hosted inference (2)

Generated from the typed local data layer. Shortlist order: verified field count → active lifecycle → source count → name. Not a ranking.

Open in selection workspace →

Open comparison builder for this use case → Pre-seeded with the top 4 shortlist candidates; these are candidates, not picks.

Create evidence brief for this use case → Markdown or JSON export of the same selection — evidence, not a recommendation.

ModelProviderLifecycleContextModalitiesSourcesFreshnessNext action
metaactive10,000,000text-in, image-in, text-out1FreshInspect the hosted-availability record and the third-party hosted pricing reference on the provider page.
deepseekactive1,000,000—3FreshReview the verified pricing reference and re-check the freshness state before projecting cost.