Skip to content
WebmasterID

Research guide

AI inference infrastructure — fields, gaps, and roadmap

Regions, cloud availability, status feeds, batching, caching, rate limits, throughput — the infrastructure dimensions a verified catalogue cares about, and which fields remain unverified today.

Last updated: 2026-05-21

What counts as infrastructure

For the purposes of the catalogue, "inference infrastructure" is the set of operational properties a builder needs to know to ship a model in production — beyond the model card itself. Stable inputs to that decision include regions, cloud availability, API surface (endpoint shape, authentication, request format), public status feeds, batching support, prompt caching, rate-limit structure, and throughput characteristics. Some are first-class fields on the catalogue today; others are documented data gaps the page reports openly.

Regions and cloud availability

A model served from a single region in one cloud is operationally different from one served across multiple regions in three clouds. The catalogue's ProviderEntity records carry a modelCatalogueUrl field that points at each vendor's region-aware model listing, but the structured per-model regions array on the ModelInfrastructure type is null today. Bedrock and Vertex model-availability matrices are the obvious next sources; they require their own verification pass.

API surface and protocols

Every verified model record carries the model parameter string used on the wire. The four major shapes in the catalogue:

  • Anthropic Messages API. POST api.anthropic.com/v1/messages with model name in the JSON body.
  • Google Gemini API. POST generativelanguage.googleapis.com/v1beta/models/<model>:generateContent with model name in the URL path.
  • DeepSeek chat completions. POST api.deepseek.com/chat/completions with model name in the JSON body.
  • Mistral chat completions. POST api.mistral.ai/v1/chat/completions with model name in the JSON body.

The differences in protocol shape — particularly model parameter position — matter for SDK adapters and for drop-in replacement scenarios. The model detail pages render a documentation-style example block for the providers whose API references are verified.

Hosted inference platforms

Some inference providers do not create their own models — they host open-weights models from other organisations and bill the developer at platform-specific rates. Groq and Together AI are the two such platforms tracked here. They are valid billing providers on a pricing record, but they are never recorded as model creators.

  • Model creator — the organisation that trained and released the model (Anthropic, Google, Meta, DeepSeek, Mistral, …). Stored on the model record's providerSlug.
  • Billing provider — the entity that invoices the developer for inference. For first-party APIs, equal to the model creator. For hosted-platform APIs, different. Stored on a pricing record's billingProviderSlug.
  • Hosted model ID — the platform-specific identifier the developer passes in API requests (e.g. Groq's meta-llama/llama-4-scout-17b-16e-instruct). Stored on a pricing record's hostedModelId.

Sprint 19 seeded the first hosted rows: Groq → Llama 4 Scout (Meta creator) and Together AI → DeepSeek V4 Pro (DeepSeek creator). Pricing decisions on those rows are made by the hosting platform — not by the model creator. We do NOT publish vendor latency / throughput / uptime claims from hosted platforms; the catalogue records only documented per-token rates, the hosted model ID, and the citation.

Sprint 20 added two further safeguards: a freshness state (Fresh / Review due / Stale / Unknown) and a volatility tag (high / medium / low / unknown) on every pricing record, and a separate hosted availability catalogue that records the stable identity claim (host × model × hosted model ID) independently from the volatile rate. WebmasterID Models does not rank hosting platforms by price. See /research/api-pricing-methodology and no-price-ranking.

Status feeds

Two patterns dominate. Anthropic and most newer providers use Statuspage, which exposes a stable JSON feed at /api/v2/status.json with a typed indicator enum. Google Cloud uses a custom incidents-feed JSON at status.cloud.google.com/incidents.json, which is an array of incident records with affected product lists. Mapping the second into our canonical ObservedStatus vocabulary requires filtering by product keyword and choosing a headline severity; see the source at lib/observers/google.ts.

Batching and caching

Every provider tracked today offers a batch API at a discounted rate (typically 50% of synchronous). Latency characteristics of the batch path are workload-dependent and provider-specific. The catalogue records the rate, not the SLA.

Prompt caching diverges sharply — Anthropic uses two TTL tiers, Google uses a one-shot write fee plus per-hour storage, DeepSeek uses an automatic cache-miss/cache-hit input distinction. The pricing-fields reference at /docs/pricing-fields documents the units; the methodology guide at /research/api-pricing-methodology explains why we keep them as separate fields.

Rate limits and throughput

Rate limits live in each vendor's account console and change with tier, region, and account history. We do not record per-tier rate-limit numbers as verified fields because they are not stable enough to be useful as a single published value; they belong in operational tooling, not in a public data catalogue.

Public vs private infrastructure fields

What providers typically publish vs keep private
Usually publicUsually private
API endpoint shapeEndpoint URL, model parameter format, payload schema.Internal routing / load-balancer topology.
PricingPer-token rates for current models; batch / cache schedule.Per-account tier pricing, enterprise discounts.
StatusVendor-reported indicator on a status page or JSON feed.Per-region health, per-instance load.
RegionsTop-level cloud availability (Bedrock, Vertex catalogues).Per-region GPU allocation, traffic shaping.
Rate limitsDefault tier limits (sometimes).Account-specific quotas, dynamic throttling thresholds.
ArchitectureModel family + tokenizer sometimes.Parameter counts, training mix, fine-tune chain, serving stack.

What we do not know

The catalogue does not know provider internals: model architecture, parameter counts, training-set size, fine-tune chain, inference-stack details. None of these are documented consistently by the providers and none can be sourced end-to-end from a primary source today. We do not estimate.

We also do not record GPU type, hardware generation, or serving topology. Those are interesting operational facts but they belong to the provider's deployment environment, not to the public-facing model.

Data roadmap

The next infrastructure-side data we expect to verify, in rough order of effort:

  1. Independent HTTP probes for more providers. Anthropic has one today; Google, DeepSeek, Mistral, OpenAI could each be onboarded with the same probe pattern (host root, no inference, no key).
  2. Durable storage for observations. The KV adapter is wired and idle; once credentials are set on production, the uptime gate at the sample threshold becomes reachable.
  3. Regional availability. Bedrock and Vertex model-availability matrices are the highest-value structured input; both have stable docs surfaces.

What this page assumes is verified

Verified today

Each item below is backed by an entry in the citation registry. Updates land via the manual verification workflow — see /docs/data-verification.

  • API endpoints for verified providers

    Anthropic Messages API (api.anthropic.com/v1/messages), Google Gemini API (generativelanguage.googleapis.com/v1beta/models/<model>:generateContent), DeepSeek chat-completions (api.deepseek.com/chat/completions), Mistral chat-completions (api.mistral.ai/v1/chat/completions). Each is recorded with the verified API reference citation.

  • Provider docs / pricing / status URLs

    Every provider entity carries primary docs, API docs, pricing docs, model catalogue, and status page URLs as verified fields where the provider publishes them.

  • Status observation pipeline

    Hourly cron writes vendor-reported and independent-probe observations for Anthropic; vendor-reported only for Google. See /research/ai-provider-status-monitoring.

Honest gaps

Data gaps

Things this page intentionally does not assert because the underlying data is not yet verified. Tracked openly so readers can calibrate.

  • Regions

    The ModelInfrastructure.regions field is null on every model record today. Provider regional availability lists exist on cloud-provider sides (Bedrock, Vertex) but require platform-specific verification we have not yet wired.

  • Average request latency

    ModelInfrastructure.avgLatencyMs is null everywhere. An independent measurement harness would be required; the only latency-shaped field we currently record is the status probe's fetch wall-clock time, which is explicitly NOT API latency.

  • Uptime percentage

    ModelInfrastructure.uptimePercent is null everywhere. See the status-monitoring policy for why we will not publish a number until we have enough durable independent observations.

  • Rate-limit ceilings

    Provider rate-limit tiers (RPM / TPM, concurrent requests, quota tiers) are not recorded as verified fields. They change frequently and are surfaced in the vendor's account console rather than in stable docs.

Continue

Related pages