What it means
Early language models only handled text. Multimodal models also accept images, and some accept audio, video and PDFs. They can describe a photo, read a scanned form, extract a table from a PDF or summarize a recorded meeting.
Most current flagship models from major providers are multimodal to some degree, but what each accepts differs, so check the documentation.
Why it matters for a business
Much business information is not plain text: invoices, receipts, site photos, screenshots and recordings. Multimodal models can read these directly, which often removes a separate OCR or transcription step.
A business example
Things to watch
-
Test accuracy on your worst real inputs, such as blurry scans.
-
Images and audio add to cost per request.
-
Check which input types each model supports.
-
Keep people reviewing low-confidence results.