Lab · structured-output
Structured extraction
Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.
Category
structured-output
Difficulty
intermediate
Estimated time
25 min
Prompts
5
When to use this
- You are wiring the model into an automation that depends on extracted fields.
- You need to confirm a candidate handles missing fields with an explicit null marker.
- You want a regression set that catches schema drift after a snapshot rotation.
Evaluation goal
Confirm the candidate produces schema-conformant JSON, marks missing fields explicitly, and never invents values to fill a schema.
Prerequisites
Export
Open the raw Markdown if you want to paste the whole set into your harness or design doc.
Prompts
5 evaluation prompts
Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.
Meeting note → attendees + decisions
Purpose
Confirm the model extracts named entities + actions without invention.
Prompt
Extract a JSON object from the source with the shape: { "attendees": string[], "decisions": { "owner": string, "action": string, "due": string | null }[], "risks": string[] } Use null for any "due" field not stated. Do not invent owners or actions. SOURCE: Attendees: Priya N., Marcus L. Date: 2025-03-04. Topic: Q2 catalogue refresh. Decisions: Marcus owns the citation backfill for the Atlas provider by 2025-03-20. Priya owns the freshness-queue spec review by 2025-03-12. No budget changes. Risks: Atlas pricing page may move under a new path; Marcus to confirm and re-run retrieval.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Valid JSON listing Priya N. and Marcus L. as attendees, with the two recorded decisions and the documented risk. All due dates are present in the source.
Failure looks like
- Invents an extra attendee.
- Returns a decision the source did not name.
- Omits the documented risk.
- Produces invalid JSON.
What to record
- JSON validity (validator pass / fail).
- Per-field accuracy against the source.
- Any invented decision or attendee.
Product spec → release + open question
Purpose
Confirm the model returns a release window string verbatim and does not promote it to a date.
Prompt
Extract a JSON object: { "feature": string, "owner": string, "release_window": string | null, "scope": string[], "out_of_scope": string[], "open_questions": string[] } Use null for any missing field. Do not invent a precise release date. SOURCE: Feature: Comparison filter chips. Owner: Daniela K. Targeted release window: 2025-Q3 (no exact date committed). Scope: filter chips render server-side; URL captures filter state; no client-side state library introduced. Out of scope: chip drag-reordering. Open question: do chips persist across navigation?Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Release window equals '2025-Q3 (no exact date committed)'. Open questions include the chip-persistence question. No invented dates.
Failure looks like
- Returns release_window as a specific calendar date.
- Drops the 'no exact date committed' qualifier.
- Invents an additional out-of-scope item.
What to record
- JSON validity.
- Whether release_window preserves the qualifier.
- Any invented field value.
Provider doc → snapshot policy
Purpose
Confirm the model returns numeric fields from the source and does not guess.
Prompt
Extract: { "snapshot_rotation_cadence": string, "deprecation_overlap_days": number | null, "notice_window_days": number | null, "notice_channels": string[] } Do not invent numeric values. Use null for any number not stated. SOURCE: Snapshot policy: production model identifiers are pinned per generation; snapshots are rotated quarterly with a 60-day deprecation overlap. Deprecation notices are published on the provider's status surface and on the model's reference page at least 30 days before retirement.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Rotation cadence 'quarterly', deprecation_overlap_days 60, notice_window_days 30, notice_channels includes the status surface and the model's reference page.
Failure looks like
- Returns deprecation_overlap_days as a value other than 60.
- Invents a notice_channel that the source did not name.
- Returns null where a number is stated.
What to record
- Numeric accuracy.
- Any invented notice channel.
Invoice → missing values must be null
Purpose
Confirm the model honours the explicit missing-value instruction.
Prompt
Extract: { "invoice_number": string, "vendor": string, "issued": string, "due": string, "line_items": { "description": string, "amount": number | null }[] } If an amount is blank in the source, the line item's amount must be null. Do not invent amounts. SOURCE: Invoice number: INV-2099-0007 (fictional). Vendor: Atlas (sample). Issued: 2025-02-14. Due: 2025-03-14. Line items: 1) "Inference compute — March 2099 — 1.2M tokens" — amount blank. 2) "Support retainer" — amount $0.00. Notes: line 1 amount intentionally left blank by sample.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Line item 1 has amount null. Line item 2 has amount 0. The fictional invoice number and vendor are preserved verbatim.
Failure looks like
- Invents an amount for line item 1.
- Treats blank as zero without the source supporting it.
- Drops a line item silently.
What to record
- Did the model return null where the source was blank?
- Did it invent an amount?
Support ticket → ambiguous fields
Purpose
Confirm the model maps explicitly-missing-or-unassigned fields to null rather than guessing.
Prompt
Extract: { "ticket_id": string, "severity": string | null, "reported_at": string, "owner": string | null, "reproduction_steps": string[] | null, "attached_logs": boolean | null } For each field, return null when the source explicitly says the field is missing or unassigned. Do not guess. SOURCE: Ticket #TKT-441-X (sample). Customer name: redacted. Severity: not set. Reported: 2025-05-01 09:15 UTC. Description: "Comparison page renders empty for some shortlist URLs." Steps to reproduce: missing. Attached logs: none. Owner: unassigned.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
severity null, owner null, reproduction_steps null, attached_logs false. ticket_id and reported_at preserved verbatim.
Failure looks like
- Returns 'unknown' or '' instead of null.
- Invents a severity.
- Returns attached_logs as null when the source says 'none'.
What to record
- Null vs string accuracy per field.
- Whether 'none' was correctly mapped to false rather than null.
Observation checklist
Tick items off in your own notes; the page does not store progress.
- Did the JSON validator pass for every prompt?
- Did the model invent any field value?
- Did the model preserve verbatim qualifiers (for example, 'no exact date committed')?
- Did the model map missing values to null rather than guessing?
- Did the model drop fields silently when uncertain?
Comparison notes
Apply when running the set across multiple candidate models.
- Pin the JSON schema bytes and reuse them across candidates — schema variance contaminates model variance.
- Capture the raw response before validation so failures can be replayed.
- Record validity rate per prompt category, not a single overall number.
How to use with the prompt-test matrix
The matrix template ships at /lab/templates/prompt-test-matrix.
- Map each prompt ID (EXT-01 … EXT-05) to a row.
- Record schema validity in the per-candidate cells (✓ / ✗ / ? with the validator output excerpt).
- Capture field-level failure modes in the rollup so the parser owner can read them.
What not to conclude
- That schema validity in this set guarantees validity for your real schema.
- That a single passing run means tool calling is reliable.
- That validity will hold across providers with different schema vocabularies.
When to rerun this set
- Your real schema changes shape.
- The provider rotates the snapshot.
- The model adds or removes a structured-output API surface.
Related playbooks and templates
Related workflows
Prompt library policy
- These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
- No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
- No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
- No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
- No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
- No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.