Lab · summarization
Summarization quality
Evaluate whether a model summarises without adding unsupported claims, omitting constraints, or inventing numbers.
Category
summarization
Difficulty
beginner
Estimated time
20 min
Prompts
5
When to use this
- You are evaluating a model that will summarise documents in your workload.
- You need to compare two candidates on how faithfully they preserve the source.
- You want a short prompt set you can rerun after a snapshot rotation.
Evaluation goal
Confirm the candidate model summarises only what the source actually says — no invented numbers, no fabricated conclusions, no dropped constraints.
Prerequisites
Export
Open the raw Markdown if you want to paste the whole set into your harness or design doc.
Prompts
5 evaluation prompts
Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.
Short summary
Purpose
Confirm the model can compress without inventing facts.
Prompt
Summarise the following text in two sentences. Do not include any fact that is not present in the source. SOURCE: The Aurora release went live on 2025-04-12 across the EU-West-1 and US-East-2 regions. Engineering tracked a 7-minute degraded-write window during the rollout but no customer-facing errors were logged. Pricing for the metered tier did not change. The next maintenance window is scheduled for 2025-05-10 with no expected downtime.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Two sentences that name the release date, regions, and that pricing did not change. The 7-minute degraded-write window may or may not appear; both are acceptable as long as nothing is invented.
Failure looks like
- Mentions an outage duration other than 7 minutes.
- Says pricing changed.
- Names a region that is not in the source.
- Returns more than three sentences.
What to record
- The output verbatim.
- Pass / fail against the acceptance rubric, with a short rationale.
- Any invented detail, even if minor.
Executive summary
Purpose
Confirm the model honours an explicit 'not stated' instruction.
Prompt
Write a three-bullet executive summary of the following text. Each bullet must reflect a fact stated in the source. If a typical executive question is not answered by the source, write "Not stated in source." instead of guessing. SOURCE: The Aurora release went live on 2025-04-12 across the EU-West-1 and US-East-2 regions. Engineering tracked a 7-minute degraded-write window during the rollout but no customer-facing errors were logged. Pricing for the metered tier did not change. The next maintenance window is scheduled for 2025-05-10 with no expected downtime.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Three bullets, each tied to a fact in the source. At least one bullet should say 'Not stated in source.' for any executive-style detail that is not present (for example, customer impact numbers).
Failure looks like
- Includes a bullet that invents a customer count.
- Returns four or more bullets.
- Drops the 'Not stated in source' instruction silently.
What to record
- Whether the 'Not stated in source' instruction was honoured.
- Bullet count.
- Any silent invention.
Bullet summary
Purpose
Surface over-interpretation behaviour.
Prompt
Return a bullet list of every concrete fact in the source. Use one bullet per fact. Do not interpret, do not infer, do not add anything that is not directly stated. SOURCE: The Aurora release went live on 2025-04-12 across the EU-West-1 and US-East-2 regions. Engineering tracked a 7-minute degraded-write window during the rollout but no customer-facing errors were logged. Pricing for the metered tier did not change. The next maintenance window is scheduled for 2025-05-10 with no expected downtime.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Bullet list strictly drawn from the source. Order may vary; content must not.
Failure looks like
- Adds a bullet about customer impact (not in source).
- Adds a bullet about cause of the degraded-write window (not stated).
- Includes a forward-looking projection (not requested).
What to record
- Bullet count.
- Bullets that introduce content not in the source.
Source-constrained summary
Purpose
Confirm the model can ground claims in literal text.
Prompt
Summarise the source. Quote at least one phrase verbatim, with surrounding quotation marks. Do not include any statement you cannot back with a quoted phrase. SOURCE: The Aurora release went live on 2025-04-12 across the EU-West-1 and US-East-2 regions. Engineering tracked a 7-minute degraded-write window during the rollout but no customer-facing errors were logged. Pricing for the metered tier did not change. The next maintenance window is scheduled for 2025-05-10 with no expected downtime.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Summary with at least one quoted phrase that appears verbatim in the source. Unsupported statements should be absent.
Failure looks like
- Returns a 'quoted' phrase that does not appear in the source.
- Includes an unsupported conclusion.
- Skips the verbatim quote instruction.
What to record
- Whether the quoted phrase actually appears verbatim.
- Any unsupported claim that snuck through.
Do-not-infer summary
Purpose
Confirm the model surfaces missing information rather than guessing.
Prompt
Summarise the source in 1–3 sentences. If the source does NOT state customer impact, you must explicitly say "Customer impact is not stated in the source." SOURCE: The Aurora release went live on 2025-04-12 across the EU-West-1 and US-East-2 regions. Engineering tracked a 7-minute degraded-write window during the rollout but no customer-facing errors were logged. Pricing for the metered tier did not change. The next maintenance window is scheduled for 2025-05-10 with no expected downtime.Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.
Expected observation
Summary that explicitly notes customer impact is not stated.
Failure looks like
- Silently omits the missing-information note.
- Invents a customer-impact assertion.
- Returns a single sentence that contradicts the source.
What to record
- Did the model surface the missing information note?
- Any overclaim about customer impact.
Observation checklist
Tick items off in your own notes; the page does not store progress.
- Did any prompt produce an invented number?
- Did any prompt produce an overconfident conclusion?
- Did the model honour explicit 'do not infer' instructions?
- Did the model preserve specific dates and region names accurately?
- Did the model silently drop instructions when they conflicted with its default style?
Comparison notes
Apply when running the set across multiple candidate models.
- Compare candidates on the same source — do not vary the source mid-suite.
- When one candidate invents a number and another does not, that is selection evidence — not a winner declaration.
- Record the verbatim outputs, not just pass/fail flags.
How to use with the prompt-test matrix
The matrix template ships at /lab/templates/prompt-test-matrix.
- Map each prompt ID (SUM-01 … SUM-05) to a row in the prompt-test-matrix template.
- Record verbatim outputs in the matrix cells — do not paraphrase.
- Use the rollup section to capture observations per candidate, not a single score.
What not to conclude
- That the candidate is 'better at summarisation' across all workloads.
- That hallucination behaviour generalises beyond this prompt set.
- That a single rerun with different temperature would replicate the result.
When to rerun this set
- The provider rotates the snapshot.
- Your summarisation workload shifts to a different document class.
- The sampling parameters you ship change.
Related playbooks and templates
Related workflows
Prompt library policy
- These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
- No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
- No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
- No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
- No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
- No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.