Claude Opus 5.5
Anthropic
Anthropic's recommended starting point for most serious work: long-running agentic coding and knowledge work, with a 1M token context window.
- Anthropic
- Vision
- Tool calling

AI models
What each model is good at, where it falls short and what to check before you build on it. Relative tiers, not benchmark scores, and a reviewed date on every page.
Featured
Anthropic
Anthropic's recommended starting point for most serious work: long-running agentic coding and knowledge work, with a 1M token context window.
Anthropic
Anthropic's balance of speed and intelligence: a strong everyday model for assistants, document work, tool calling and coding.
Anthropic
Anthropic's fastest and lowest-cost Claude model, with near-frontier intelligence for high-volume and real-time work.
Google's most advanced Gemini model for reasoning, software engineering and agent work, reading text, images, audio, video and PDFs. Available as a preview.
Google's most capable Flash model, stable since September 2026, for agents, software engineering and enterprise workflows with full multimodal input.
OpenAI
The middle model of the GPT-6 family, positioned for complex coding and agentic workflows at a mid price level.
OpenAI
OpenAI's most efficient model for focused, high-volume tasks, with vision, tool calling and structured outputs at the lowest price level in the family.
Directory
45 of 45 models
Anthropic
The top tier of the Claude family, for the most demanding reasoning and long-horizon agent work, priced above Opus.
Anthropic
Anthropic's recommended starting point for most serious work: long-running agentic coding and knowledge work, with a 1M token context window.
Anthropic
Anthropic's balance of speed and intelligence: a strong everyday model for assistants, document work, tool calling and coding.
Anthropic
Anthropic's fastest and lowest-cost Claude model, with near-frontier intelligence for high-volume and real-time work.
Google's most advanced Gemini model for reasoning, software engineering and agent work, reading text, images, audio, video and PDFs. Available as a preview.
Google's most capable Flash model, stable since September 2026, for agents, software engineering and enterprise workflows with full multimodal input.
Google's lowest-cost current Gemini model for high-throughput work such as sub-agent tasks and document parsing.
Google's open-weight model family under Apache 2.0, in sizes from phone-friendly to 31B, with image input and function calling.
Google's current embedding model, multimodal: it embeds text, images, video, audio and PDFs for search and retrieval.
Google's Gemini-native image generation and editing models, known as Nano Banana 2 and Nano Banana Pro, which replace Imagen in the Gemini API.
Google's video generation model, creating short video with native audio from text and images. Available as a preview, with fast and lite versions.
Deepgram
Deepgram's speech-to-text models: Nova-3 for general transcription in 60+ languages, and Flux for voice agents with built-in end-of-turn detection.
ElevenLabs
ElevenLabs text-to-speech: the expressive Eleven v3 in 70+ languages, a real-time conversational version, and the low-latency Flash models, with voice cloning.
Stability AI
Stability AI's open-weight image models, in Large, Medium and Flash versions, which can be self-hosted or used through the Stability API.
Black Forest Labs
Black Forest Labs' FLUX.2 image models: API-only Pro, Flex and Max versions, plus open-weight Dev and Klein versions under different licenses.
Midjourney
Midjourney's image generation service, known for its visual style. V8.1 is the default since June 2026 and V8.2 followed in July. There is no public developer API.
Runway
Runway's video models: Gen-4.5 for text-to-video and image-to-video, and Aleph 2.0 for editing existing video, with a developer API.
Meta
An open-weight Llama 4 model with image input and a very long context window, for teams that want to host a capable model themselves.
Meta
The larger Llama 4 open-weight model, with image input and a 1 million token context window, for self-hosted assistants and analysis.
Meta
An older, text-only open-weight Llama model with tool use and a 128K token window, still widely hosted and well understood.
Meta
Meta's newer proprietary model family, offered through the Meta Model API, with text, image, video and PDF input, tool calling and a 1 million token window.
Mistral AI
Mistral's open-weight general-purpose flagship under Apache 2.0, with image input, tool calling, structured outputs and a 256K token window.
Mistral AI
Mistral's efficient open-weight model under Apache 2.0, one hybrid model for instructions, reasoning and coding, with tool calling and a 256K window.
Mistral AI
A newer open-weight Mistral model for agent and coding work, with image input, tool calling, structured outputs and a 256K window.
Mistral AI
Mistral's low-latency coding model for code completion and generation, with fill-in-the-middle support, tool calling and a 128K window.
Mistral AI
Mistral's document OCR model, turning pages into structured output with bounding boxes, block labels and confidence scores.
OpenAI
OpenAI's most capable model, built for the hardest end-to-end work: complex reasoning, coding, computer use and research.
OpenAI
The middle model of the GPT-6 family, positioned for complex coding and agentic workflows at a mid price level.
OpenAI
OpenAI's most efficient model for focused, high-volume tasks, with vision, tool calling and structured outputs at the lowest price level in the family.
OpenAI
An older OpenAI model described as its smartest non-reasoning model, strong at following instructions and calling tools, with a very long context window.
OpenAI
A compact, low-cost older OpenAI model for focused tasks, with vision, tool calling and structured outputs and a 128K token window.
OpenAI
OpenAI's most capable embedding model for search and retrieval, with 3,072-dimension vectors and support for English and other languages.
OpenAI
OpenAI's efficient, lowest-cost embedding model, with 1,536-dimension vectors for search, retrieval and similarity at scale.
OpenAI
OpenAI's current speech-to-text model, replacing Whisper and the GPT-4o transcribe models, with streaming and keyword or language hints.
OpenAI
OpenAI's current image generation and editing models: Sunburst for the highest quality and Flare for fast everyday images.
OpenAI
OpenAI's current text-to-speech model for turning text into natural spoken audio, at a low price level.
DeepSeek
DeepSeek's default, best-value model: open weights under MIT, image input, tool calling, JSON output, optional thinking and a 1 million token window.
DeepSeek
DeepSeek's model for demanding reasoning and agent work, with thinking effort levels, tool calling, JSON output and a 1 million token window.
Alibaba Cloud
Alibaba's current Qwen flagship for long autonomous coding and professional work, with image and video input, tool calling, JSON Schema output and a 1 million token window.
Alibaba Cloud
Qwen's open-weight coding models under Apache 2.0, from the efficient Qwen3-Coder-Next to the large 480B model, with tool use and long context.
xAI
xAI's current flagship for coding, agent tasks and knowledge work, with image input, tool calling, structured outputs, reasoning effort levels and a 500K window.
Cohere
Cohere's enterprise flagship for multimodal, multilingual agent tasks, with 48 languages, tool calling, JSON output and open weights under Apache 2.0.
Cohere
Cohere's multimodal embedding model for enterprise search, embedding text, images and mixed documents, with flexible vector sizes and long inputs.
Cohere
Cohere's multilingual rerank models, which sort search results by relevance, in a best-quality Pro version and a low-latency Fast version.
Voyage AI by MongoDB
Voyage AI's current embedding series, now part of MongoDB, with large, standard and lite models that share one embedding space.
Nothing matches those filters. or ask us.
Browse
Models that think through multi-step problems before answering: analysis, planning, math, complex documents and agent work.
All-round language models for writing, summarizing, support, extraction and most everyday business tasks.
Models that write, review and explain code, and power coding assistants and developer tools.
Models that read images, scans, screenshots and sometimes audio or video alongside text.
Models that turn text into vectors for search, retrieval, clustering and recommendations.
Speech-to-text and text-to-speech models for transcription, call analysis and voice assistants.
Models that generate or edit images and video from text and reference images.
Models whose weights you can download and run on your own servers or a cloud of your choice.
Model picker
Six quick questions give you a category and two or three models to test first. It is a starting point, not a verdict.
How to read this directory
Each model page shows eight relative tiers from one to five: reasoning, coding, vision, speed, price level, context size, tool calling and structured output. They compare models with each other at the time of review. They are our judgment from provider documentation and our own use, not benchmark results.
Every page also lists what the model is good at, where it is not the right choice, the business tasks we would use it for and what to check before you commit. Prices and limits change often, so each page carries a reviewed date and a link to the provider documentation.
To choose for a real project, we test two or three candidates on your own examples. That is what our AI model evaluation service does, and it usually takes days, not weeks. See also LLM integration, MCP servers and the problems we solve with these models.
Keep exploring
FAQ
Have a question that is not here? Ask us directly.
There is no single best model. The right one depends on the task, volume, speed, budget and data rules. Use the model picker on this page for a starting point, then test two or three candidates on your own examples. We do this as part of AI model evaluation.
Benchmarks change often and rarely match business tasks. Relative tiers give a fair first impression. Real decisions should come from tests on your own data.
Every model page shows the date we last checked it against the provider documentation. Prices and limits change often, so always confirm them with the provider before you commit.
Most production systems use more than one: a fast, low-cost model for simple steps and a stronger model for the hard ones. Building with a thin layer between your code and the provider also makes switching easier later. See LLM integration.
Start a project
Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.
Loading the search index...
Search is not available right now. Try the HTML sitemap.
No matches. Try a shorter word, or browse the resources.
Type at least two letters. Press / anywhere to open search.
We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.
Necessary
Remembers your theme and this choice. Always on.
Analytics
Google Analytics 4, to count visits and see which pages help. No advertising.
Maps and embeds
Loads the Google map on the contact page and other outside content.