Skip to content
WebmasterID

Lab · intermediate

Structured output testing

How to validate JSON mode, structured output, and tool calls against your real schema before depending on the model in a pipeline.

Difficulty

intermediate

Estimated time

30 min

Output

2 Markdown artifacts

A planning recipe, not a safety validation.

Schema validity is workload-specific. The playbook does not validate your schema for compliance or safety, and does not declare which model is best for structured generation.

Goal

Confirm a candidate model produces schema-conformant output reliably enough for the parser downstream.

When to use this

  • You are wiring the model into an automation that depends on a fixed schema.
  • You have read the structured-output lesson and need to validate the verified feature claim against your own schema.
  • You need to compare two candidates on schema reliability without ranking them on a generic benchmark.

Prerequisites

Test setup

  • Choose the structured-generation surface that matches your candidate (JSON mode, structured output, tool calling).
  • Pin your schema in a file — every run uses the exact same schema bytes.
  • Pick a strict validator (Ajv, Pydantic, zod, your own) and run validation on every response.
  • Capture the raw response before validation so you can replay failures.

Minimum test set

  • 10–20 prompts covering the schema's common, edge, and adversarial shapes.
  • A prompt that asks the model to refuse — verify refusal still emits valid schema if your contract requires it.
  • A prompt with conflicting instructions — verify the model picks one shape rather than emitting hybrid output.

Prompt variants

Vary one dimension at a time so you can attribute behaviour.

  • Same prompt, schema field reordering, to surface ordering-dependent failures.
  • Same prompt, schema with optional fields removed, to confirm the model does not invent values.
  • Same prompt, low and high temperature, to measure schema reliability under sampling.

Observations to record

Record observations, not scores. The evidence brief stays auditable.

  • Schema validity rate per candidate, per prompt category.
  • Field-level failure modes (missing fields, extra fields, type mismatch, enum violation).
  • Latency overhead of structured generation vs free-form generation for the same prompt.
  • How the candidate behaves when the schema is malformed (does it refuse, hallucinate, or hang?).

Failure modes to watch

  • Model emits valid JSON that fails enum validation — likely tokenizer or alignment issue.
  • Model adds extra fields the schema disallows — common with newer snapshots.
  • Model truncates output mid-structure when max_tokens is hit.
  • Tool-call surface returns arguments as strings when the schema requires numbers.

Stop conditions

Knowing when to stop is part of the test.

  • Stop and re-scope if schema validity drops below your acceptance threshold across happy-path prompts.
  • Stop and document if the same schema fails on one candidate and passes on another — that is selection evidence, not a recommendation.

Outputs

  • A Markdown evidence brief listing per-prompt validity for each candidate.
  • The exact schema bytes used in the suite, attached for reproducibility.

Weak test vs stronger test

Weak test

  • Eyeball one JSON response and declare the integration ready.
  • Skip running every response through a strict validator.
  • Vary the schema, the prompt, and the sampling parameters all at once and call the result a single test.

Stronger test

  • Pin the exact schema bytes and reuse them across every run.
  • Run a strict validator (Ajv / Pydantic / zod) on every response before logging.
  • Vary one dimension at a time so failures attribute to a single change.
  • Capture raw responses pre-validation so failures replay deterministically.

Observation rubric

Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.

DimensionWhat to look forWhat to record
Schema validityWhether the response passes a strict validator against your real schema.Pass / fail per prompt with the raw response captured alongside.
Field-level failure modeMissing fields, extra fields, type mismatches, enum violations.Per-prompt failure category and an example excerpt.
Latency overheadLatency delta between structured and free-form generation for the same prompt.Median delta per prompt category in milliseconds.
Behaviour under malformed schemaDoes the model refuse, hallucinate, or hang when handed a malformed schema?Per-case observation with the malformed input captured.

Record this in your brief

Brief content that keeps the playbook output reusable for a reviewer.

  • Attach the schema bytes so the brief is reproducible.
  • Record validity rate per prompt category, not as a single percentage.
  • List field-level failure modes the parser would have to absorb.
  • Capture latency overhead so the reviewer knows what structured generation costs.

Related templates

Related workflows

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.