Skip to content

Glossary

What is a multimodal model?

A multimodal model is an AI model that can take in more than one kind of input, such as text with images, audio or video.

What it means

Early language models only handled text. Multimodal models also accept images, and some accept audio, video and PDFs. They can describe a photo, read a scanned form, extract a table from a PDF or summarize a recorded meeting.

Most current flagship models from major providers are multimodal to some degree, but what each accepts differs, so check the documentation.

Why it matters for a business

Much business information is not plain text: invoices, receipts, site photos, screenshots and recordings. Multimodal models can read these directly, which often removes a separate OCR or transcription step.

A business example

Things to watch

  • Test accuracy on your worst real inputs, such as blurry scans.

  • Images and audio add to cost per request.

  • Check which input types each model supports.

  • Keep people reviewing low-confidence results.

Keep exploring

FAQ

Questions about multimodal model

Have a question that is not here? Ask us directly.

Start a project

Tell us what you want to build. We will show you a faster path.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.