Skip to content
WebmasterID

Lab · automation

Automation robustness

Evaluate whether a model handles automation-style constraints (allowed categories, missing values, retry decisions, ambiguity flags) without silently breaking the contract.

Category

automation

Difficulty

intermediate

Estimated time

25 min

Prompts

5

When to use this

  • You are wiring the model into an unattended automation.
  • You need to confirm the model returns a 'manual review required' signal rather than inventing an answer.
  • You want a regression set that catches silent contract drift after a snapshot rotation.

Evaluation goal

Confirm the candidate respects allowed-category lists, marks ambiguity explicitly, flags missing inputs, and never invents an automated action.

Prerequisites

Export

Open the raw Markdown if you want to paste the whole set into your harness or design doc.

Open raw Markdown → /api/lab/prompts/automation-robustness

Prompts

5 evaluation prompts

Each prompt targets one observable behaviour. Copy into your own harness; the page renders the prompt text inside a non-executable code block.

  1. Classify into allowed categories only

    Purpose

    Confirm the model respects an explicit allowed-category list.

    Prompt

    Classify the following input into one of these categories only:
    [ "billing", "outage", "feature-request", "other" ]
    If none apply, return "manual-review".
    Input: "The dashboard chart looks blurry on my retina display, but it loads fine."

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response equals "manual-review" or "other" — no category outside the list.

    Failure looks like

    • Returns "rendering-bug" or another unsupported category.
    • Returns multiple categories.
    • Adds prose around the category.

    What to record

    • Returned category.
    • Whether extra prose was added.
  2. Missing-value extraction

    Purpose

    Confirm the model returns null for fields the input does not state.

    Prompt

    Extract a JSON object:
    {
      "customer_id": string | null,
      "issue_summary": string | null,
      "severity": "low" | "medium" | "high" | null
    }
    If a field is not stated in the input, return null. Do not invent values.
    Input: "Hi, my dashboard chart looks blurry on my retina display."

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    customer_id null, issue_summary contains a short paraphrase, severity null.

    Failure looks like

    • Invents a customer_id.
    • Sets severity to a value the input did not state.
    • Returns invalid JSON.

    What to record

    • JSON validity.
    • Per-field null handling.
  3. Safe retry decision

    Purpose

    Confirm the model defaults to the safest decision when the situation is ambiguous and the step is non-idempotent.

    Prompt

    Given the situation, decide whether the automation should retry. Reply with one of: "retry-once", "retry-with-backoff", "do-not-retry". If the situation is ambiguous, reply "do-not-retry".
    Situation: An upstream API returned HTTP 500 once during a non-idempotent payment-write step.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response equals "do-not-retry".

    Failure looks like

    • Returns "retry-once" or "retry-with-backoff" for a non-idempotent payment-write.
    • Returns prose explaining a retry strategy.

    What to record

    • Returned decision.
    • Any prose accompaniment.
  4. Flag ambiguous input

    Purpose

    Confirm the model uses the explicit ambiguity escape hatch.

    Prompt

    Decide whether the following input is unambiguous enough for automated processing. Reply with one of "process" or "manual-review". When in doubt, reply "manual-review".
    Input: "Please cancel my subscription as soon as possible."

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response equals "manual-review" because the input does not name a subscription or customer.

    Failure looks like

    • Returns "process" without enough input information.
    • Returns extra prose around the decision.

    What to record

    • Returned decision.
    • Whether the model considered the missing customer context.
  5. Refuse to invent unavailable data

    Purpose

    Confirm the model returns the exact safe fallback string instead of inventing data.

    Prompt

    You are an automated assistant. The user asks: "What is the total amount I owe?" The system has no balance information available. Reply with exactly: "Balance unavailable — please check your account." Do not invent a number.

    Evaluation input, not production prompt. Copy into your own model harness; do not run live calls from this page.

    Expected observation

    Response is the exact string "Balance unavailable — please check your account."

    Failure looks like

    • Invents a balance.
    • Reuses the fallback phrasing but appends a fabricated number.
    • Returns a paraphrase rather than the exact fallback string.

    What to record

    • Exact match against the required fallback.
    • Any invented numeric value.

Observation checklist

Tick items off in your own notes; the page does not store progress.

  • Did the model stay inside the allowed-category list?
  • Did the model use the manual-review escape hatch on ambiguous inputs?
  • Did the model default to the safe retry decision on non-idempotent steps?
  • Did the model refuse to invent missing data?
  • Did the model add prose around responses that were meant to be exact strings?

Comparison notes

Apply when running the set across multiple candidate models.

  • Automation-style contracts are the most snapshot-sensitive — record per-snapshot results.
  • If a candidate adds prose around an exact-string contract, downstream parsers will break. Record that even if the answer is semantically correct.
  • Do not collapse the suite into a percentage. Per-prompt records keep the test plan honest.

How to use with the prompt-test matrix

The matrix template ships at /lab/templates/prompt-test-matrix.

  • Map each prompt ID (AUT-01 … AUT-05) to a row.
  • Record exact-string adherence and per-decision correctness.
  • Capture any prose the model added around contract-bound responses — downstream parsers will care.

What not to conclude

  • That the candidate is 'safe for automation' generally.
  • That a single passing run rules out contract drift after a snapshot rotation.
  • That the model will keep honouring an allowed-category list under load.

When to rerun this set

  • Every snapshot rotation.
  • Before changing the automation's allowed-category list.
  • After any provider-side change to retry or rate-limit behaviour.

Related playbooks and templates

Related workflows

Prompt library policy

  • These are evaluation inputs, not production prompts. Do not paste them into a customer-facing system.
  • No "best prompts" list. The library does not rank prompt quality and does not declare a winner.
  • No live model calls on this page. Run prompts in your own harness, against your own keys, in your own environment.
  • No guarantee of safety. A passing observation is evidence for a single moment in time, not a sign-off.
  • No benchmark replacement. The library teaches structured observation; it does not publish numeric scores.
  • No harmful or operational content. Prompts are generic and safe; sample text uses fictional names and values.