Skip to content
WebmasterID

Learn · model fundamentals

Structured output, JSON mode, and tool use

The difference between structured output, JSON mode, and tool/function calling — and what is currently verified in the catalogue.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

Three different concepts often used interchangeably

Providers describe similar but distinct capabilities under overlapping names. Inspecting verified fields means knowing what each one actually does:

  • JSON mode — the model is constrained to produce syntactically valid JSON. The shape is up to the prompt; the API only enforces "this is JSON".
  • Structured output — the model is constrained to produce JSON that conforms to a schema you pass with the request. The provider enforces the schema, not just the syntax.
  • Tool / function calling — the model can emit calls to predefined tools or functions with structured arguments, often used to delegate work back to your code.

These three capabilities have different API surfaces, different latency profiles, and different failure modes. Treating them as one is a recurring source of integration bugs.

What is currently verified in the catalogue

The catalogue records the features field on each model as a verified set of capability tags. Capability verification is gated on the provider explicitly naming the feature in their documentation — not on third-party tutorials or blog posts. A model that supports a capability without it being verified renders the canonical unverified-data label.

Why not infer capability from marketing text

"Supports structured output" in a marketing blurb can mean any of the three things above. The catalogue therefore requires the provider's official docs to name the API surface (e.g. response_format, tool_choice) before the feature lands as a verified field. Until then, the gap is explicit.

What to verify before integrating structured generation

  • The provider documents the API surface (response_format, tool_choice, JSON schema, function-call schema) explicitly.
  • Your schema is supported (some providers restrict OpenAPI / JSON Schema features).
  • Latency overhead of structured generation is acceptable in your environment.
  • Failure modes are predictable — malformed schema, refusal, truncated output — and your code handles each one.
  • If you use tool calling, the catalogue's features field names tool calling explicitly for the model.

Common mistakes

  • Treating JSON mode as structured output.

    JSON mode guarantees syntax, not shape. A model can return valid JSON that fails your schema validation.

  • Building tool calls before confirming the model lists the feature.

    Tool calling is a separate API surface. A model with structured output may not support tool calling.

  • Using one provider's schema vocabulary across providers.

    JSON Schema subsets and OpenAPI extensions vary. Always confirm against the target provider's docs.

Apply this workflow

Apply this workflow

Data gaps to watch

Structured-generation capabilities are the most rapidly evolving fields in the catalogue. Treat the unverified-data label as "the catalogue has not yet retrieved a primary-source citation that names this capability for this model" rather than "this capability does not exist."

Related pages

Sources and freshness

Feature citations age the fastest of any field in the catalogue. The reverification queue prioritises capability fields for re-check whenever a provider publishes a new snapshot.

Teaching example

Illustrative — not a recommendation.

Situation: A pipeline depends on the model returning JSON that matches a fixed schema. The team needs to confirm the candidate honours an explicit schema rather than just returning valid-looking JSON.

Decision to make: Which structured-generation surface (JSON mode, structured output, tool calling) does the candidate actually verify, and against what schema features?

Verified fields that matter:

  • Features field (catalogue capability tag)
  • API surface name (response_format, tool_choice, etc.)
  • Schema vocabulary supported
  • Source citation for the capability

Weak vs better approach

Weak approach

  • Treat 'JSON mode' as equivalent to schema-conformant output.
  • Skip schema validation on the response.
  • Use one provider's schema vocabulary against another's API.
  • Assume tool calling works because JSON mode is listed.

Better approach

  • Distinguish JSON mode, structured output, and tool calling explicitly.
  • Validate every response against your real schema with a strict validator.
  • Confirm the schema vocabulary the candidate accepts (subset of JSON Schema, etc.).
  • Test each capability in isolation against your real prompts.

Why better: The three capabilities have different API surfaces and different failure modes. The better approach catches schema-shape drift that JSON-syntax checks miss.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Structured-output inspection note

## Structured-output inspection (illustrative)
Candidate: <slug>
Capability listed (verified): <json-mode / structured-output / tool-calling>
API surface: <response_format / tools / function_call>
Schema bytes used in trial: <hash or filename>

Trial results (5 prompts):
- P-01 schema-valid: yes
- P-02 schema-valid: no — extra field <name> outside schema
- P-03 schema-valid: yes
- P-04 schema-valid: yes
- P-05 schema-valid: no — type mismatch on <field>

Action: surface schema failures in the brief; do not infer integration-ready.

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Distinguish JSON mode, structured output, and tool calling.

  2. Step 2

    Apply in /compare/build →

    Compare verified feature tags across candidates.

  3. Step 3

    Verify in /coverage →

    Audit per-provider feature coverage and citation density.

  4. Step 4

    Test in /lab →

    Run the structured-output testing playbook + prompt set.

Review before moving on

  • I have named which structured-generation surface I depend on.
  • I have a fixed schema and a strict validator.
  • I have run the structured-extraction prompt set against the candidate.
  • I have NOT assumed tool calling works because JSON mode is listed.
  • Schema failures are recorded in the brief, not collapsed to a percentage.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Asserting which model produces the most reliable structured output — that depends on schema, prompt, and workload.
  • Asserting that JSON mode and structured output and tool use mean the same thing — they do not.
  • Replacing your own schema-driven evaluation work.