Lab · intermediate
Multimodal input testing
How to test image, audio, video, and PDF input channels against your real assets — never against marketing copy.
Difficulty
intermediate
Estimated time
30 min
Output
2 Markdown artifacts
A planning recipe, not a safety validation.
Modality support is workload-specific. The playbook does not declare which model is best for any modality and does not certify accessibility, accuracy, or safety of the model's outputs.
Goal
Confirm that a candidate model handles the modality your workload actually sends, on assets that resemble your production traffic.
When to use this
- Your application sends image, audio, video, or PDF input.
- You need to confirm the model accepts your real asset format, not just the demo format.
- You need to understand failure modes when an asset is malformed or partial.
Prerequisites
Test setup
- Verify the modality channel is present as a verified field on the candidate model — if not, the test confirms whether the gap is real.
- Confirm the API surface accepts the asset encoding you plan to send (base64, URL, multi-part, signed-link).
- Cap asset sizes at your real workload's 95th percentile, not the API's documented maximum.
- Capture asset hashes so failures are reproducible.
Minimum test set
- 5–10 happy-path assets representative of your typical traffic.
- 2–3 edge-case assets (low resolution, low bitrate, scanned PDFs, multi-page documents).
- 2–3 adversarial assets (malformed, partial, mismatched MIME type).
Prompt variants
Vary one dimension at a time so you can attribute behaviour.
- Same asset, different question framings — surfaces prompt-vs-asset attribution failures.
- Same asset, with and without an OCR pre-pass when relevant.
- Same asset, fed via two different transports (URL vs base64) to surface transport-dependent failures.
Observations to record
Record observations, not scores. The evidence brief stays auditable.
- Pass / fail per asset class against your acceptance rubric.
- Whether the model refuses, hallucinates, or returns a useful error on adversarial input.
- Latency vs the same prompt without the multimodal asset.
- Cost per asset broken down by token unit if pricing is asymmetric across modalities.
Failure modes to watch
- The model passes happy-path assets but silently falls back to text-only when the image fails to decode.
- Audio transcription drifts on accented speech the model was not trained on.
- PDF parsing returns body text but drops headers, footers, or table structure.
- The model invents content for blank pages or silent audio.
Stop conditions
Knowing when to stop is part of the test.
- Stop and re-scope if the candidate fails on assets that match your real traffic distribution.
- Stop and document if the model silently degrades modality — silent fallbacks are integration hazards.
Outputs
- A Markdown evidence brief with per-asset-class pass rates.
- A list of asset hashes attached for replay.
Weak test vs stronger test
Weak test
- Send one demo-quality image and assume real assets behave the same.
- Skip the silent-fallback test.
- Skip asset-size validation.
Stronger test
- Send 5–10 happy-path assets matched to your real workload distribution.
- Include adversarial assets (malformed, low resolution, partial).
- Probe explicitly for silent fallback (model returns text-only without erroring).
- Capture asset hashes so failures replay.
Observation rubric
Observations, not scores. Each row names what to look at and what to record — there is no aggregate number.
| Dimension | What to look for | What to record |
|---|---|---|
| Happy-path accuracy | Whether the model handles representative assets against your rubric. | Pass / fail per asset with a rationale, plus the asset hash. |
| Adversarial behaviour | Refusal, hallucination, or useful error on malformed / partial assets. | Per-asset category outcome and a short note on safety implications. |
| Silent fallback | Whether the model silently drops to text-only when the modality fails. | Yes / no per case with the response that triggered the observation. |
| Latency vs text-only baseline | Latency delta vs the same prompt without the multimodal asset. | Median delta per asset class. |
Record this in your brief
Brief content that keeps the playbook output reusable for a reviewer.
- Record per-asset-class pass rates instead of a single accuracy number.
- List asset hashes alongside outcomes so failures are reproducible.
- Flag any silent fallback explicitly — silent fallback is an integration hazard.
- Capture the latency delta vs text-only.
Related templates
Related workflows
- Multimodal use case
- Comparison builder
- Create a decision brief after testing
- Review sources and freshness before signing off.
What the lab does not promise
- No production readiness guarantee. A passing playbook is evidence, not approval.
- No compliance or regulatory certification. Verification is not certification.
- No safety validation. Templates and playbooks are planning tools, not safety reviews.
- No model ranking. The lab does not score candidates against each other.
- No benchmark replacement. The lab teaches your own testing discipline; it does not publish synthesized benchmark numbers.