AI DevelopmentAnalysis5 min readPublished September 30, 2026

A shared embedding space changes the indexing and query decision

Cohere Embed 5: Pro and Fast Embedding Models Compared

Cohere Embed 5 comes in Pro and Fast versions with 128K context. How the two compare on price, speed and use cases, and when to switch embedding models.

DA
Digital Applied Team
Research and practical guidance
CoverageSeptember 30, 2026

Embed 5’s most useful design choice is that Pro and Fast share an embedding space. A team can evaluate higher-quality corpus indexing and faster query embedding without maintaining two incompatible indexes within that family. The choice still needs testing on the actual documents and questions; shared space is not a guarantee of equal retrieval quality.

Editorial note: Prepared October 1 as a September 30, 2026 dispatch, using the dated announcements cited below. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Pro and Fast can work togetherCohere recommends indexing with Pro and querying with Fast.
  2. 02
    Text and image prices differUse the appropriate input unit when estimating corpus cost.
  3. 03
    Treat older-index migration separatelyShared space within Embed 5 does not establish compatibility with older model vectors.

01 — The evidenceThe two variants and published prices

Cohere’s September 30 launch lists Embed 5 Pro at $0.12 per million text tokens and Fast at $0.08; image tokens cost $0.40 per million for either. The model documentation and dated release note describe a shared space, 128K context, text/image inputs and more than 100 languages. The launch announcement supplies the prices; these are vendor rates, not a measured workload bill.

Sources: Cohere Embed 5 launch and model documentation, September 30, 2026. Deployment-specific charges and storage are outside this table.
PropertyProFast
Text / 1M tokens$0.12$0.08
Image / 1M tokens$0.40$0.40
Context128K tokens128K tokens
Vendor positioningQuality-focused indexing and retrieval.Latency-sensitive query traffic.
Vector dimensions256–2048, six supported sizes.Same supported sizes.

The supported dimensions are 256, 512, 768, 1024, 1536 and 2048, with float, int8 and binary outputs. Select a consistent representation for the index and query path. A vector database may impose its own requirements on dimensions, distance metric and supported formats.

02 — Practical implicationsWhat a shared space changes

An embedding is a numerical representation used to compare a query with stored material. When two models produce compatible representations, the index built with one can be queried with the other as documented. Cohere recommends Pro for indexing and Fast for queries, which separates an infrequent corpus operation from repeated interactive traffic.

Corpus
Prepare the searchable material
Indexing choice

Test Pro on the document collection, including difficult layouts and languages.

Stored vectors
Query
Represent the user’s question
Serving choice

Compare Fast and Pro against the same index and relevance judgments.

Interactive path
Answer
Use the retrieved evidence
Generation choice

Check that the final answer cites and uses the correct retrieved material.

End-to-end result

Do not generalize compatibility beyond the stated family. An existing index made with another embedding model should be treated as a separate migration. Matching vector length alone does not prove that distances have the same meaning. Keep the old index until the replacement has passed retrieval and application checks.

03 — Practical implicationsLonger inputs do not eliminate document design

A 128K context window allows larger inputs, but the best retrieval unit may still be a section, page or other coherent chunk. A whole document can contain several unrelated answers. If the stored representation blends them together, the system may retrieve a broadly relevant document without locating the passage needed for a precise response.

Compare chunking strategies on actual questions. Include queries that require a table row, an exception clause or information near the end of a long document. Preserve identifiers that let the application link a retrieved vector back to the right source location. A high similarity score is less useful when the user cannot inspect the evidence.

For visually structured documents, compare parsed text with the supported image or mixed-input path. Use the same relevance judgments and keep parsing failures visible. The model cannot recover information that a preprocessing step removed before embedding.

04 — Practical implicationsRun a small parallel-index evaluation

Digital Applied migration procedure; no hands-on Embed 5 benchmark is claimed.
StepEvidence to keep
Freeze inputsA versioned document sample and representative query set.
Build a candidateModel, dimensions, format, chunking and source identifiers.
Judge retrievalRelevant results and important misses for each query.
Measure servingLatency, throughput and actual usage under the expected load.
Test rollbackA verified route back to the previous index and query model.

Use human relevance judgments where the answer matters, and inspect failures by category. An average can hide poor results for one language or one document format. Keep permissions and metadata filters in the evaluation too: a relevant document that the user is not allowed to access is a failed application result.

Illustratively, embedding 100 million text tokens at the listed rates costs $12 with Pro or $8 with Fast, excluding every other charge. That $4 difference may be small beside the cost of parsing, storing and reviewing the corpus. For frequent query traffic, latency and repeated usage can matter more. Recalculate from your actual volumes.

05 — Practical implicationsChoose the pair that improves useful retrieval

The model-routing guide explains accepted-result economics; the agent-memory guide places retrieval in a larger system. Keep the access boundary fixed while testing. Our AI transformation service supports those controlled evaluations.

Next step

Compare query quality before replacing the existing index

Use the shared space to test a practical Pro/Fast pairing, then choose from relevance, latency and complete operating cost. Preserve the old index until the new pipeline handles both useful answers and difficult misses correctly.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

Cohere documents a shared Embed 5 space and recommends that pairing, subject to consistent configuration.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source