Skip to content
WebmasterID

Lab

AI Usage Lab

Practical playbooks for testing AI models before production use — prompt tests, structured-output checks, long-context trials, multimodal trials, automation-risk reviews, and regression checks. Paste-ready Markdown templates are included. Learn → Apply → Verify → Test.

The lab extends the curriculum from Learn → Apply → Verify into Test. Each playbook is a testing recipe you run yourself before integrating a model; each template is a paste-ready Markdown planning document you adapt to your workload.

Lab workflow

  1. Step 1

    Define task →

    Name the workload, the acceptance rubric, and the data gaps.

  2. Step 2

    Build test set →

    Pick representative prompts (or assets) from real traffic and pin a fixed schema.

  3. Step 3

    Run model trials →

    Execute the playbook against each candidate model with parameters held constant.

  4. Step 4

    Record evidence →

    Roll observations into an evidence brief and store the canary suite for regression checks.

Playbooks

6 testing playbooks

Each playbook walks one testing dimension — prompt behaviour, structured output, long-context, multimodal, automation, regression — and ends with a Markdown evidence brief you can attach to /briefs/build.

Prompts

Evaluation prompt library

Six prompt sets — summarisation, structured extraction, long-context recall, instruction following, refusal boundary, automation robustness. Evaluation inputs you run in your own harness, not production prompts.

All prompt sets →

New here? Start at /lab/prompt-testing-basics or read the evaluation guide.

Templates

3 paste-ready templates

Generic Markdown planning documents you adapt per workload. Every template is exportable via /api/lab/templates/<slug>.

All templates →

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.