Skip to content

AI model by Cohere

Cohere Rerank 4 for business

Cohere's multilingual rerank models, which sort search results by relevance, in a best-quality Pro version and a low-latency Fast version.

  • Reviewed on September 24, 2026
  • Cohere

Capability tiersRelative, not benchmarks

Reasoning Very low
Coding Very low
Vision Very low
Speed High
Price level Low
Context size Medium
Tool calling Very low
Structured output Very low

Key facts about Cohere Rerank 4

Model ID at review
rerank-v4.0-pro, rerank-v4.0-fast
Provider
Cohere
Main category
Embeddings
Open weights
No, available as a hosted service
Last reviewed
September 24, 2026

Fit

Where it fits and where it does not

Good at

  • Improving the order of search results.

  • Raising answer quality in RAG assistants.

  • Working on top of any existing search.

  • Multilingual queries and documents.

  • A choice between quality and speed.

Not the right choice for

  • Searching a whole collection on its own.

  • Generating answers.

  • Documents longer than its 32K token window per item.

  • Self-hosting.

Use cases

Business use cases we would use it for

Better RAG answers

Reranking retrieved passages before a model answers. See RAG knowledge assistants.

Site and help search

Sorting help center results by true relevance. See internal help desk.

Product search

Improving the order of product results for natural questions.

Our notes

When we would choose it

Rerank 4 is Cohere's current rerank line, superseding Rerank 3.5. It comes in two versions: rerank-v4.0-pro for the best quality and rerank-v4.0-fast for low latency. Both are multilingual, with a 32K token context.

A reranker takes the top results from a first search, whether keyword, embedding or both, and sorts them by how well each one answers the query. It is a small addition that often gives one of the largest improvements in a knowledge assistant, because the language model can only answer from what reaches it.

Rerank works with any first-stage search. You do not need to change your embedding model or your database to add it. That makes it one of the cheapest upgrades we suggest for existing assistants that sometimes miss the right passage.

We would measure the gain directly: take a set of real questions, record how often the right passage appears in the top three results with and without reranking, and weigh that against the added delay and cost.

For live search where every millisecond counts, start with the Fast version. For back-office research tools, the Pro version is usually worth the extra time. See Cohere Embed and embedding models.

Before you commit

Things to check before you commit

  • Latency

    Rerank adds a step. Use the Fast version for live search.

  • Gain

    Measure how often the right passage reaches the top three, before and after.

  • Input size

    The context is 32K tokens. Chunk long documents.

  • Older versions

    Rerank 4 supersedes v3.5. Test the switch.

Alternatives

Models to compare it with

Cohere Embed v4

Cohere's multimodal embedding model for enterprise search, embedding text, images and mixed documents, with flexible vector sizes and long inputs.

Voyage 4

Voyage AI's current embedding series, now part of MongoDB, with large, standard and lite models that share one embedding space.

text-embedding-3-small

OpenAI's efficient, lowest-cost embedding model, with 1,536-dimension vectors for search, retrieval and similarity at scale.

Keep exploring

FAQ

Questions people ask us

Have a question that is not here? Ask us directly.

Start a project

Not sure which model fits? Ask us to evaluate your use case.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.