Skip to content

AI model category

Vision and multimodal AI models

For invoices, forms, photos, screenshots and other inputs that are not plain text.

  • 25 models

Multimodal models accept images, and sometimes audio or video, together with text. In business systems they read scanned invoices, forms and receipts, describe site photos, check screenshots and extract tables from PDFs. Many general purpose models now include vision, so this category overlaps with others.

For documents, a multimodal model often replaces a separate OCR step. For high volumes of the same document type, a dedicated OCR model can still be cheaper.

Vision and multimodal

All vision and multimodal ai models

Reasoning

Claude Fable 5.1

Anthropic

The top tier of the Claude family, for the most demanding reasoning and long-horizon agent work, priced above Opus.

  • Anthropic
  • Vision
  • Tool calling
Reasoning

Claude Opus 5.5

Anthropic

Anthropic's recommended starting point for most serious work: long-running agentic coding and knowledge work, with a 1M token context window.

  • Anthropic
  • Vision
  • Tool calling
General purpose

Claude Sonnet 5

Anthropic

Anthropic's balance of speed and intelligence: a strong everyday model for assistants, document work, tool calling and coding.

  • Anthropic
  • Vision
  • Tool calling
General purpose

Claude Haiku 4.5

Anthropic

Anthropic's fastest and lowest-cost Claude model, with near-frontier intelligence for high-volume and real-time work.

  • Anthropic
  • Vision
  • Tool calling
Reasoning

Gemini 3.1 Pro

Google

Google's most advanced Gemini model for reasoning, software engineering and agent work, reading text, images, audio, video and PDFs. Available as a preview.

  • Google
  • Vision
  • Tool calling
General purpose

Gemini 3.8 Flash

Google

Google's most capable Flash model, stable since September 2026, for agents, software engineering and enterprise workflows with full multimodal input.

  • Google
  • Vision
  • Tool calling
General purpose

Gemini 3.5 Flash-Lite

Google

Google's lowest-cost current Gemini model for high-throughput work such as sub-agent tasks and document parsing.

  • Google
  • Vision
  • Tool calling
Open weights

Gemma 4

Google

Google's open-weight model family under Apache 2.0, in sizes from phone-friendly to 31B, with image input and function calling.

  • Google
  • Open weights
  • Vision
  • Tool calling
Embeddings

Gemini Embedding 2

Google

Google's current embedding model, multimodal: it embeds text, images, video, audio and PDFs for search and retrieval.

  • Google
  • Vision
Open weights

Llama 4 Scout

Meta

An open-weight Llama 4 model with image input and a very long context window, for teams that want to host a capable model themselves.

  • Meta
  • Open weights
  • Vision
  • Tool calling
Open weights

Llama 4 Maverick

Meta

The larger Llama 4 open-weight model, with image input and a 1 million token context window, for self-hosted assistants and analysis.

  • Meta
  • Open weights
  • Vision
  • Tool calling
General purpose

Meta Muse Spark

Meta

Meta's newer proprietary model family, offered through the Meta Model API, with text, image, video and PDF input, tool calling and a 1 million token window.

  • Meta
  • Vision
  • Tool calling
General purpose

Mistral Large 3

Mistral AI

Mistral's open-weight general-purpose flagship under Apache 2.0, with image input, tool calling, structured outputs and a 256K token window.

  • Mistral AI
  • Open weights
  • Vision
  • Tool calling
Coding

Mistral Medium 3.5

Mistral AI

A newer open-weight Mistral model for agent and coding work, with image input, tool calling, structured outputs and a 256K window.

  • Mistral AI
  • Open weights
  • Vision
  • Tool calling
Vision and multimodal

Mistral OCR 4.1

Mistral AI

Mistral's document OCR model, turning pages into structured output with bounding boxes, block labels and confidence scores.

  • Mistral AI
  • Vision
Reasoning

GPT-6 Astra

OpenAI

OpenAI's most capable model, built for the hardest end-to-end work: complex reasoning, coding, computer use and research.

  • OpenAI
  • Vision
  • Tool calling
General purpose

GPT-6 Sol

OpenAI

The middle model of the GPT-6 family, positioned for complex coding and agentic workflows at a mid price level.

  • OpenAI
  • Vision
  • Tool calling
General purpose

GPT-6 Luna

OpenAI

OpenAI's most efficient model for focused, high-volume tasks, with vision, tool calling and structured outputs at the lowest price level in the family.

  • OpenAI
  • Vision
  • Tool calling
General purpose

GPT-4.1

OpenAI

An older OpenAI model described as its smartest non-reasoning model, strong at following instructions and calling tools, with a very long context window.

  • OpenAI
  • Vision
  • Tool calling
General purpose

GPT-4o mini

OpenAI

A compact, low-cost older OpenAI model for focused tasks, with vision, tool calling and structured outputs and a 128K token window.

  • OpenAI
  • Vision
  • Tool calling
General purpose

DeepSeek V4.1 Flash

DeepSeek

DeepSeek's default, best-value model: open weights under MIT, image input, tool calling, JSON output, optional thinking and a 1 million token window.

  • DeepSeek
  • Open weights
  • Vision
  • Tool calling
Reasoning

Qwen3.8-Max

Alibaba Cloud

Alibaba's current Qwen flagship for long autonomous coding and professional work, with image and video input, tool calling, JSON Schema output and a 1 million token window.

  • Alibaba Cloud
  • Vision
  • Tool calling
Reasoning

Grok 4.7

xAI

xAI's current flagship for coding, agent tasks and knowledge work, with image input, tool calling, structured outputs, reasoning effort levels and a 500K window.

  • xAI
  • Vision
  • Tool calling
General purpose

Cohere Command A+

Cohere

Cohere's enterprise flagship for multimodal, multilingual agent tasks, with 48 languages, tool calling, JSON output and open weights under Apache 2.0.

  • Cohere
  • Open weights
  • Vision
  • Tool calling
Embeddings

Cohere Embed v4

Cohere

Cohere's multimodal embedding model for enterprise search, embedding text, images and mixed documents, with flexible vector sizes and long inputs.

  • Cohere
  • Vision

How to choose

How to choose in this category

Collect fifty real documents or images, including poor scans and unusual layouts, and measure field accuracy for each candidate. Check image size limits, price per page and how the model handles handwriting and tables. Keep a person in the loop for low-confidence results.

Other categories

Browse other categories

Keep exploring

Start a project

Tell us what you want to build. We will show you a faster path.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.