Lab · intermediate
Structured output testing
How to validate JSON mode, structured output, and tool calls against your real schema before depending on the model in a pipeline.
Difficulty
intermediate
Estimated time
30 min
Output
2 Markdown artifacts
A planning recipe, not a safety validation.
Schema validity is workload-specific. The playbook does not validate your schema for compliance or safety, and does not declare which model is best for structured generation.
Goal
Confirm a candidate model produces schema-conformant output reliably enough for the parser downstream.
When to use this
- You are wiring the model into an automation that depends on a fixed schema.
- You have read the structured-output lesson and need to validate the verified feature claim against your own schema.
- You need to compare two candidates on schema reliability without ranking them on a generic benchmark.
Prerequisites
Test setup
- Choose the structured-generation surface that matches your candidate (JSON mode, structured output, tool calling).
- Pin your schema in a file — every run uses the exact same schema bytes.
- Pick a strict validator (Ajv, Pydantic, zod, your own) and run validation on every response.
- Capture the raw response before validation so you can replay failures.
Minimum test set
- 10–20 prompts covering the schema's common, edge, and adversarial shapes.
- A prompt that asks the model to refuse — verify refusal still emits valid schema if your contract requires it.
- A prompt with conflicting instructions — verify the model picks one shape rather than emitting hybrid output.
Prompt variants
Vary one dimension at a time so you can attribute behaviour.
- Same prompt, schema field reordering, to surface ordering-dependent failures.
- Same prompt, schema with optional fields removed, to confirm the model does not invent values.
- Same prompt, low and high temperature, to measure schema reliability under sampling.
Observations to record
Record observations, not scores. The evidence brief stays auditable.
- Schema validity rate per candidate, per prompt category.
- Field-level failure modes (missing fields, extra fields, type mismatch, enum violation).
- Latency overhead of structured generation vs free-form generation for the same prompt.
- How the candidate behaves when the schema is malformed (does it refuse, hallucinate, or hang?).
Failure modes to watch
- Model emits valid JSON that fails enum validation — likely tokenizer or alignment issue.
- Model adds extra fields the schema disallows — common with newer snapshots.
- Model truncates output mid-structure when max_tokens is hit.
- Tool-call surface returns arguments as strings when the schema requires numbers.
Stop conditions
Knowing when to stop is part of the test.
- Stop and re-scope if schema validity drops below your acceptance threshold across happy-path prompts.
- Stop and document if the same schema fails on one candidate and passes on another — that is selection evidence, not a recommendation.
Outputs
- A Markdown evidence brief listing per-prompt validity for each candidate.
- The exact schema bytes used in the suite, attached for reproducibility.
Weak test vs stronger test
Weak test
- Eyeball one JSON response and declare the integration ready.
- Skip running every response through a strict validator.
- Vary the schema, the prompt, and the sampling parameters all at once and call the result a single test.
Stronger test
- Pin the exact schema bytes and reuse them across every run.
- Run a strict validator (Ajv / Pydantic / zod) on every response before logging.
- Vary one dimension at a time so failures attribute to a single change.
- Capture raw responses pre-validation so failures replay deterministically.
Observation rubric
Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.
| Dimension | What to look for | What to record |
|---|---|---|
| Schema validity | Whether the response passes a strict validator against your real schema. | Pass / fail per prompt with the raw response captured alongside. |
| Field-level failure mode | Missing fields, extra fields, type mismatches, enum violations. | Per-prompt failure category and an example excerpt. |
| Latency overhead | Latency delta between structured and free-form generation for the same prompt. | Median delta per prompt category in milliseconds. |
| Behaviour under malformed schema | Does the model refuse, hallucinate, or hang when handed a malformed schema? | Per-case observation with the malformed input captured. |
Record this in your brief
Brief content that keeps the playbook output reusable for a reviewer.
- Attach the schema bytes so the brief is reproducible.
- Record validity rate per prompt category, not as a single percentage.
- List field-level failure modes the parser would have to absorb.
- Capture latency overhead so the reviewer knows what structured generation costs.
Related templates
Related workflows
- Comparison builder
- Decision brief builder
- Create a decision brief after testing
- Review sources and freshness before signing off.
What the lab does not promise
- No production readiness guarantee. A passing playbook is evidence, not approval.
- No compliance or regulatory certification. Verification is not certification.
- No safety validation. Templates and playbooks are planning tools, not safety reviews.
- No model ranking. The lab does not score candidates against each other.
- No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.