Skip to content
WebmasterID

Learn · model fundamentals

Multimodal input: image, audio, video, PDF

How the catalogue records which models accept image, audio, video, or PDF input — and why marketing copy is not enough to assume support.

Last reviewed 2026-05-24. Lesson copy is reviewed when the underlying catalogue policy changes — not on a fixed cadence.

What counts as a modality channel

Modality is the kind of content a model can accept as input or produce as output. The catalogue records modality as a verified field on each model, enumerating channels such as text-in, image-in, audio-in, video-in, and text-out. A channel only appears when the provider's documentation explicitly names it as supported.

Why modality must be source-backed

Provider marketing copy often describes models as "multimodal" without enumerating the channels. A reader who treats that claim as verified ends up shipping an integration that calls an image endpoint the model does not support, or a PDF parser that silently drops to text-only mode. The catalogue therefore records modality only when the provider's official docs name each channel — and renders the canonical unverified-data label for the rest.

What to verify on modality

  • The model's modality field is verified, not just described as 'multimodal' in marketing.
  • Each input channel you need (image, audio, video, PDF) is enumerated explicitly.
  • Each output channel (text, structured output, function calls) is enumerated separately.
  • The citation for the modality field points at official provider docs, not a blog post.
  • Any unverified channel renders the unverified-data label — that is a question for your tests, not a missing feature.

How to inspect modality in the catalogue

The selection workspace can filter the shortlist by modality. Pass ?modality=image-in in the URL to narrow to models with a verified image-input channel. The model record page renders the full channel list under the modality field, with the citation alongside.

Common mistakes

  • Assuming 'multimodal' means all modalities are supported.

    Different providers use the word for very different channel sets. Always read the enumerated list.

  • Treating PDF support as the same as image support.

    Some models parse PDFs as text, some render pages as images, some support both. The channel name matters.

  • Ignoring output channels.

    A model that accepts images may still only output text. Confirm both directions before integrating.

  • Inferring modality from a single demo video.

    Demos use curated inputs. Real workloads expose channel limits the demo did not.

Apply this workflow

Apply this workflow

Data gaps to watch

When a model's modality field renders the unverified-data label, the catalogue has not yet recorded a primary-source citation that enumerates the channels. Confirm externally before integrating, or queue the source via /reverification.

Related pages

Sources and freshness

Modality citations age slowly compared to pricing, but they do change when providers add new input channels. Check /sources for the underlying URLs and /reverification for anything due for re-check.

Teaching example

Illustrative — not a recommendation.

Situation: An application needs to extract data from scanned invoices uploaded as PDFs. The team has shortlisted three candidates described in marketing as 'multimodal'.

Decision to make: Which candidates have a verified PDF input channel, and which silently fall back to text-only?

Verified fields that matter:

  • Verified modality channels (text-in, image-in, audio-in, video-in, etc.)
  • Explicit PDF support (if enumerated by the provider)
  • Input transport accepted (URL, base64, multipart)
  • Documented size or page caps

Weak vs better approach

Weak approach

  • Trust the word 'multimodal' on the marketing page.
  • Assume PDF support if image support is listed.
  • Skip the silent-fallback test.
  • Skip asset-size validation.

Better approach

  • Confirm modality channels are an enumerated, verified field.
  • Test each candidate against a representative sample of your assets.
  • Probe for silent fallbacks (model returns text-only without erroring).
  • Cap asset sizes at your real workload's 95th percentile.

Why better: Silent modality fallback is the most expensive multimodal failure mode because it produces plausible-looking text instead of an error. The better approach surfaces the fallback before integration.

Example artifact

Illustrative example — not a recommendation. Substitute your own values when you run the workflow.

Modality verification note

## Modality verification (illustrative)
Workload: scanned-invoice extraction (PDFs, 2–8 pages)

| Candidate | image-in | pdf-in | source URL | retrievedAt |
| --- | --- | --- | --- | --- |
| <model-A> | verified | verified | <docs> | <date> |
| <model-B> | verified | unverified-data | <docs> | <date> |

Test plan additions:
- For <model-B>, run a 5-PDF sample to confirm whether PDF support is implicit.
- Record any silent fallback to text-only.

Substitute your real values when you walk the workflow. The catalogue never generates this artifact for you.

Concept → workflow bridge

  1. Step 1

    Learn the concept →

    Confirm the input channels the catalogue actually verifies.

  2. Step 2

    Apply in /select →

    Filter the catalogue to models with verified image input.

  3. Step 3

    Verify in /sources →

    Read the modality citation behind each model's verified channels.

  4. Step 4

    Test in /lab →

    Run the multimodal input testing playbook on your real assets.

Review before moving on

  • Each candidate's modality channels are an enumerated, verified field.
  • I have a test plan for assets matching my real traffic distribution.
  • I will probe each candidate for silent text-only fallback.
  • Asset-size caps match my workload's 95th percentile.
  • I have NOT inferred PDF support from generic 'multimodal' copy.

Caution: No persistence — the checklist resets on every visit. Capture progress in your own notes.

What this lesson does not teach

  • Asserting which model is best for multimodal workloads — the catalogue records modality channels, not quality rankings.
  • Claiming a model supports a modality the catalogue has not verified — unverified channels render the unverified-data label.
  • Replacing your own workload-specific testing of multimodal inputs.