Skip to content
WebmasterID

Lab · intermediate

Multimodal input testing

How to test image, audio, video, and PDF input channels against your real assets — never against marketing copy.

Difficulty

intermediate

Estimated time

30 min

Output

2 Markdown artifacts

A planning recipe, not a safety validation.

Modality support is workload-specific. The playbook does not declare which model is best for any modality and does not certify accessibility, accuracy, or safety of the model's outputs.

Goal

Confirm that a candidate model handles the modality your workload actually sends, on assets that resemble your production traffic.

When to use this

  • Your application sends image, audio, video, or PDF input.
  • You need to confirm the model accepts your real asset format, not just the demo format.
  • You need to understand failure modes when an asset is malformed or partial.

Prerequisites

Test setup

  • Verify the modality channel is present as a verified field on the candidate model — if not, the test confirms whether the gap is real.
  • Confirm the API surface accepts the asset encoding you plan to send (base64, URL, multi-part, signed-link).
  • Cap asset sizes at your real workload's 95th percentile, not the API's documented maximum.
  • Capture asset hashes so failures are reproducible.

Minimum test set

  • 5–10 happy-path assets representative of your typical traffic.
  • 2–3 edge-case assets (low resolution, low bitrate, scanned PDFs, multi-page documents).
  • 2–3 adversarial assets (malformed, partial, mismatched MIME type).

Prompt variants

Vary one dimension at a time so you can attribute behaviour.

  • Same asset, different question framings — surfaces prompt-vs-asset attribution failures.
  • Same asset, with and without an OCR pre-pass when relevant.
  • Same asset, fed via two different transports (URL vs base64) to surface transport-dependent failures.

Observations to record

Record observations, not scores. The evidence brief stays auditable.

  • Pass / fail per asset class against your acceptance rubric.
  • Whether the model refuses, hallucinates, or returns a useful error on adversarial input.
  • Latency vs the same prompt without the multimodal asset.
  • Cost per asset broken down by token unit if pricing is asymmetric across modalities.

Failure modes to watch

  • The model passes happy-path assets but silently falls back to text-only when the image fails to decode.
  • Audio transcription drifts on accented speech the model was not trained on.
  • PDF parsing returns body text but drops headers, footers, or table structure.
  • The model invents content for blank pages or silent audio.

Stop conditions

Knowing when to stop is part of the test.

  • Stop and re-scope if the candidate fails on assets that match your real traffic distribution.
  • Stop and document if the model silently degrades modality — silent fallbacks are integration hazards.

Outputs

  • A Markdown evidence brief with per-asset-class pass rates.
  • A list of asset hashes attached for replay.

Weak test vs stronger test

Weak test

  • Send one demo-quality image and assume real assets behave the same.
  • Skip the silent-fallback test.
  • Skip asset-size validation.

Stronger test

  • Send 5–10 happy-path assets matched to your real workload distribution.
  • Include adversarial assets (malformed, low resolution, partial).
  • Probe explicitly for silent fallback (model returns text-only without erroring).
  • Capture asset hashes so failures replay.

Observation rubric

Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.

DimensionWhat to look forWhat to record
Happy-path accuracyWhether the model handles representative assets against your rubric.Pass / fail per asset with a rationale, plus the asset hash.
Adversarial behaviourRefusal, hallucination, or useful error on malformed / partial assets.Per-asset category outcome and a short note on safety implications.
Silent fallbackWhether the model silently drops to text-only when the modality fails.Yes / no per case with the response that triggered the observation.
Latency vs text-only baselineLatency delta vs the same prompt without the multimodal asset.Median delta per asset class.

Record this in your brief

Brief content that keeps the playbook output reusable for a reviewer.

  • Record per-asset-class pass rates instead of a single accuracy number.
  • List asset hashes alongside outcomes so failures are reproducible.
  • Flag any silent fallback explicitly — silent fallback is an integration hazard.
  • Capture the latency delta vs text-only.

Related templates

Related workflows

What the lab does not promise

  • No production readiness guarantee. A passing playbook is evidence, not approval.
  • No compliance or regulatory certification. Verification is not certification.
  • No safety validation. Templates and playbooks are planning tools, not safety reviews.
  • No model ranking. The lab does not score candidates against each other.
  • No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.