Claude Fable 5.1
Anthropic
The top tier of the Claude family, for the most demanding reasoning and long-horizon agent work, priced above Opus.
- Anthropic
- Vision
- Tool calling
AI model category
For invoices, forms, photos, screenshots and other inputs that are not plain text.
Multimodal models accept images, and sometimes audio or video, together with text. In business systems they read scanned invoices, forms and receipts, describe site photos, check screenshots and extract tables from PDFs. Many general purpose models now include vision, so this category overlaps with others.
For documents, a multimodal model often replaces a separate OCR step. For high volumes of the same document type, a dedicated OCR model can still be cheaper.
Vision and multimodal
Anthropic
The top tier of the Claude family, for the most demanding reasoning and long-horizon agent work, priced above Opus.
Anthropic
Anthropic's recommended starting point for most serious work: long-running agentic coding and knowledge work, with a 1M token context window.
Anthropic
Anthropic's balance of speed and intelligence: a strong everyday model for assistants, document work, tool calling and coding.
Anthropic
Anthropic's fastest and lowest-cost Claude model, with near-frontier intelligence for high-volume and real-time work.
Google's most advanced Gemini model for reasoning, software engineering and agent work, reading text, images, audio, video and PDFs. Available as a preview.
Google's most capable Flash model, stable since September 2026, for agents, software engineering and enterprise workflows with full multimodal input.
Google's lowest-cost current Gemini model for high-throughput work such as sub-agent tasks and document parsing.
Google's open-weight model family under Apache 2.0, in sizes from phone-friendly to 31B, with image input and function calling.
Google's current embedding model, multimodal: it embeds text, images, video, audio and PDFs for search and retrieval.
Meta
An open-weight Llama 4 model with image input and a very long context window, for teams that want to host a capable model themselves.
Meta
The larger Llama 4 open-weight model, with image input and a 1 million token context window, for self-hosted assistants and analysis.
Meta
Meta's newer proprietary model family, offered through the Meta Model API, with text, image, video and PDF input, tool calling and a 1 million token window.
Mistral AI
Mistral's open-weight general-purpose flagship under Apache 2.0, with image input, tool calling, structured outputs and a 256K token window.
Mistral AI
A newer open-weight Mistral model for agent and coding work, with image input, tool calling, structured outputs and a 256K window.
Mistral AI
Mistral's document OCR model, turning pages into structured output with bounding boxes, block labels and confidence scores.
OpenAI
OpenAI's most capable model, built for the hardest end-to-end work: complex reasoning, coding, computer use and research.
OpenAI
The middle model of the GPT-6 family, positioned for complex coding and agentic workflows at a mid price level.
OpenAI
OpenAI's most efficient model for focused, high-volume tasks, with vision, tool calling and structured outputs at the lowest price level in the family.
OpenAI
An older OpenAI model described as its smartest non-reasoning model, strong at following instructions and calling tools, with a very long context window.
OpenAI
A compact, low-cost older OpenAI model for focused tasks, with vision, tool calling and structured outputs and a 128K token window.
DeepSeek
DeepSeek's default, best-value model: open weights under MIT, image input, tool calling, JSON output, optional thinking and a 1 million token window.
Alibaba Cloud
Alibaba's current Qwen flagship for long autonomous coding and professional work, with image and video input, tool calling, JSON Schema output and a 1 million token window.
xAI
xAI's current flagship for coding, agent tasks and knowledge work, with image input, tool calling, structured outputs, reasoning effort levels and a 500K window.
Cohere
Cohere's enterprise flagship for multimodal, multilingual agent tasks, with 48 languages, tool calling, JSON output and open weights under Apache 2.0.
Cohere
Cohere's multimodal embedding model for enterprise search, embedding text, images and mixed documents, with flexible vector sizes and long inputs.
How to choose
Collect fifty real documents or images, including poor scans and unusual layouts, and measure field accuracy for each candidate. Check image size limits, price per page and how the model handles handwriting and tables. Keep a person in the loop for low-confidence results.
Other categories
Keep exploring
Start a project
Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.
Loading the search index...
Search is not available right now. Try the HTML sitemap.
No matches. Try a shorter word, or browse the resources.
Type at least two letters. Press / anywhere to open search.
We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.
Necessary
Remembers your theme and this choice. Always on.
Analytics
Google Analytics 4, to count visits and see which pages help. No advertising.
Maps and embeds
Loads the Google map on the contact page and other outside content.