Skip to content
WebmasterID

Exercise · intermediate

Create a decision evidence brief

Use the decision brief builder to generate a paste-ready evidence pack from your shortlist, then export it in Markdown.

Estimated time: 10 minutes. Exercises produce evidence artifacts, never model recommendations.

Goal

Produce a Markdown evidence brief that captures verified fields, data gaps, source trails, and freshness — ready for the next reviewer.

Prerequisites

  • Read /learn/how-to-choose-ai-model and /learn/testing-ai-models so the brief framing is clear.
  • Have a shortlist or comparison URL from the earlier exercises.

Step-by-step exercise

  1. Open the decision brief builder

    Visit /briefs/build and select 2–4 candidate models, optionally with a use-case filter.

    Open /briefs/build →

    Expected outcome: The page renders the evidence summary, verified fields, data gaps, source trail, and freshness notes.

  2. Read the evidence summary

    Confirm the brief lists every verified field with its citation, and every data gap as an explicit gap row.

    Open /briefs/build →

    Expected outcome: You can name how many evidence fields are verified vs how many are gaps.

  3. Export the brief in Markdown

    Open /api/briefs/decision?models=<slug1>,<slug2>&useCase=<slug> to get the Markdown export.

    Open /api/briefs/decision →

    Expected outcome: You have a Markdown file you can paste into a design doc, ticket, or PR description.

  4. Read the example brief

    Compare your brief to the published example to confirm structure parity.

    Open /examples/decision-brief →

    Expected outcome: You can describe how the catalogue's brief differs from a model recommendation — same fields, no verdict.

Completion checklist

Completion checklist

  • Your brief lists verified fields with citations.
  • Your brief lists data gaps explicitly.
  • Your brief carries a freshness note.
  • You did NOT add a 'recommendation' section to the brief.

Evidence artifact

At the end of this exercise you should have: A Markdown brief from /api/briefs/decision that any reviewer can read independently.

Paste the artifact into your design doc, ticket, or PR description. The catalogue's role ends with the artifact; the workload-specific testing is yours.

Related workflow routes

Exercise does not recommend a model — external testing still required.

Common mistake

Adding a 'Recommendation' section to the exported brief. The catalogue produces evidence; recommendations belong in your team's decision doc, not in the brief.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Example artifact

## Decision brief excerpt (illustrative Markdown)

### Evidence summary
Candidates: <slug-A>, <slug-B>, <slug-C>
Use case: <use-case-slug>
Verified fields: context window, max output, modality, lifecycle, pricing references

### Data gaps
- <slug-B> max output: unverified-data label
- <slug-C> hosted pricing: stale (re-verify before quoting)

### Source trail
- <slug-A> · provider docs · retrievedAt <date>
- <slug-B> · provider docs · retrievedAt <date>
- <slug-C> · provider docs · retrievedAt <date>

Note: brief intentionally contains no recommendation.

Substitute your real shortlist, slugs, dates, and values when you walk the exercise.

Repeat this exercise when

  • The shortlist changes.
  • A pricing reference is re-verified.
  • A snapshot rotation invalidates a field in the brief.
  • The reviewer asks for an updated cut for a meeting.

Review before moving on

  • Brief lists verified fields with citations.
  • Data gaps are explicit, not hidden.
  • Freshness note is present.
  • No 'recommendation' or 'winner' section.

Caution: No persistence — checklist resets on every visit. Capture progress in your own notes.

What this exercise does not produce

  • A model recommendation. The exercise routes you through evidence — you decide.
  • A score or grade for any model. The catalogue does not score.
  • A substitute for external prompt, latency, rate-limit, cost, or compliance tests.