Skip to content
WebmasterID

Lab · structured-output

Structured extraction

Evaluate whether a model extracts fields into a requested structure without inventing missing values or breaking schema constraints.

Category

structured-output

Difficulty

intermediate

Estimated time

25 min

Prompts

5

When to use this

  • You are wiring the model into an automation that depends on extracted fields.
  • You need to confirm a candidate handles missing fields with an explicit null marker.
  • You want a regression set that catches schema drift after a snapshot rotation.

Evaluation goal

Confirm the candidate produces schema-conformant JSON, marks missing fields explicitly, and never invents values to fill a schema.

Prerequisites

Export

Open the raw Markdown if you want to paste the whole set into your harness or design doc.

Open raw Markdown → /api/lab/prompts/structured-extraction

Prompts

5 evaluation prompts

Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.

  1. Meeting note → attendees + decisions

    Purpose

    Confirm the model extracts named entities + actions without invention.

    Prompt

    Extract a JSON object from the source with the shape:
    {
      "attendees": string[],
      "decisions": { "owner": string, "action": string, "due": string | null }[],
      "risks": string[]
    }
    Use null for any "due" field not stated. Do not invent owners or actions.
    
    SOURCE:
    Attendees: Priya N., Marcus L. Date: 2025-03-04. Topic: Q2 catalogue refresh. Decisions: Marcus owns the citation backfill for the Atlas provider by 2025-03-20. Priya owns the freshness-queue spec review by 2025-03-12. No budget changes. Risks: Atlas pricing page may move under a new path; Marcus to confirm and re-run retrieval.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Valid JSON listing Priya N. and Marcus L. as attendees, with the two recorded decisions and the documented risk. All due dates are present in the source.

    Failure looks like

    • Invents an extra attendee.
    • Returns a decision the source did not name.
    • Omits the documented risk.
    • Produces invalid JSON.

    What to record

    • JSON validity (validator pass / fail).
    • Per-field accuracy against the source.
    • Any invented decision or attendee.
  2. Product spec → release + open question

    Purpose

    Confirm the model returns a release window string verbatim and does not promote it to a date.

    Prompt

    Extract a JSON object:
    {
      "feature": string,
      "owner": string,
      "release_window": string | null,
      "scope": string[],
      "out_of_scope": string[],
      "open_questions": string[]
    }
    Use null for any missing field. Do not invent a precise release date.
    
    SOURCE:
    Feature: Comparison filter chips. Owner: Daniela K. Targeted release window: 2025-Q3 (no exact date committed). Scope: filter chips render server-side; URL captures filter state; no client-side state library introduced. Out of scope: chip drag-reordering. Open question: do chips persist across navigation?

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Release window equals '2025-Q3 (no exact date committed)'. Open questions include the chip-persistence question. No invented dates.

    Failure looks like

    • Returns release_window as a specific calendar date.
    • Drops the 'no exact date committed' qualifier.
    • Invents an additional out-of-scope item.

    What to record

    • JSON validity.
    • Whether release_window preserves the qualifier.
    • Any invented field value.
  3. Provider doc → snapshot policy

    Purpose

    Confirm the model returns numeric fields from the source and does not guess.

    Prompt

    Extract:
    {
      "snapshot_rotation_cadence": string,
      "deprecation_overlap_days": number | null,
      "notice_window_days": number | null,
      "notice_channels": string[]
    }
    Do not invent numeric values. Use null for any number not stated.
    
    SOURCE:
    Snapshot policy: production model identifiers are pinned per generation; snapshots are rotated quarterly with a 60-day deprecation overlap. Deprecation notices are published on the provider's status surface and on the model's reference page at least 30 days before retirement.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Rotation cadence 'quarterly', deprecation_overlap_days 60, notice_window_days 30, notice_channels includes the status surface and the model's reference page.

    Failure looks like

    • Returns deprecation_overlap_days as a value other than 60.
    • Invents a notice_channel that the source did not name.
    • Returns null where a number is stated.

    What to record

    • Numeric accuracy.
    • Any invented notice channel.
  4. Invoice → missing values must be null

    Purpose

    Confirm the model honours the explicit missing-value instruction.

    Prompt

    Extract:
    {
      "invoice_number": string,
      "vendor": string,
      "issued": string,
      "due": string,
      "line_items": { "description": string, "amount": number | null }[]
    }
    If an amount is blank in the source, the line item's amount must be null. Do not invent amounts.
    
    SOURCE:
    Invoice number: INV-2099-0007 (fictional). Vendor: Atlas (sample). Issued: 2025-02-14. Due: 2025-03-14. Line items: 1) "Inference compute — March 2099 — 1.2M tokens" — amount blank. 2) "Support retainer" — amount $0.00. Notes: line 1 amount intentionally left blank by sample.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Line item 1 has amount null. Line item 2 has amount 0. The fictional invoice number and vendor are preserved verbatim.

    Failure looks like

    • Invents an amount for line item 1.
    • Treats blank as zero without the source supporting it.
    • Drops a line item silently.

    What to record

    • Did the model return null where the source was blank?
    • Did it invent an amount?
  5. Support ticket → ambiguous fields

    Purpose

    Confirm the model maps explicitly-missing-or-unassigned fields to null rather than guessing.

    Prompt

    Extract:
    {
      "ticket_id": string,
      "severity": string | null,
      "reported_at": string,
      "owner": string | null,
      "reproduction_steps": string[] | null,
      "attached_logs": boolean | null
    }
    For each field, return null when the source explicitly says the field is missing or unassigned. Do not guess.
    
    SOURCE:
    Ticket #TKT-441-X (sample). Customer name: redacted. Severity: not set. Reported: 2025-05-01 09:15 UTC. Description: "Comparison page renders empty for some shortlist URLs." Steps to reproduce: missing. Attached logs: none. Owner: unassigned.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    severity null, owner null, reproduction_steps null, attached_logs false. ticket_id and reported_at preserved verbatim.

    Failure looks like

    • Returns 'unknown' or '' instead of null.
    • Invents a severity.
    • Returns attached_logs as null when the source says 'none'.

    What to record

    • Null vs string accuracy per field.
    • Whether 'none' was correctly mapped to false rather than null.

Observation checklist

Tick items off in your own notes; the page does not store progress.

  • Did the JSON validator pass for every prompt?
  • Did the model invent any field value?
  • Did the model preserve verbatim qualifiers (for example, 'no exact date committed')?
  • Did the model map missing values to null rather than guessing?
  • Did the model drop fields silently when uncertain?

Comparison notes

Apply when running the set across multiple candidate models.

  • Pin the JSON schema bytes and reuse them across candidates — schema variance contaminates model variance.
  • Capture the raw response before validation so failures can be replayed.
  • Record validity rate per prompt category, not a single overall number.

How to use with the prompt-test matrix

The matrix template ships at /lab/templates/prompt-test-matrix.

  • Map each prompt ID (EXT-01 … EXT-05) to a row.
  • Record schema validity in the per-candidate cells (✓ / ✗ / ? with the validator output excerpt).
  • Capture field-level failure modes in the rollup so the parser owner can read them.

What not to conclude

  • That schema validity in this set guarantees validity for your real schema.
  • That a single passing run means tool calling is reliable.
  • That validity will hold across providers with different schema vocabularies.

When to rerun this set

  • Your real schema changes shape.
  • The provider rotates the snapshot.
  • The model adds or removes a structured-output API surface.

Related playbooks and templates

Related workflows

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.