Embed 5’s most useful design choice is that Pro and Fast share an embedding space. A team can evaluate higher-quality corpus indexing and faster query embedding without maintaining two incompatible indexes within that family. The choice still needs testing on the actual documents and questions; shared space is not a guarantee of equal retrieval quality.
Editorial note: Prepared October 1 as a September 30, 2026 dispatch, using the dated announcements cited below. Later product developments are outside this article’s scope.
- 01Pro and Fast can work togetherCohere recommends indexing with Pro and querying with Fast.
- 02Text and image prices differUse the appropriate input unit when estimating corpus cost.
- 03Treat older-index migration separatelyShared space within Embed 5 does not establish compatibility with older model vectors.
01 — The evidenceThe two variants and published prices
Cohere’s September 30 launch lists Embed 5 Pro at $0.12 per million text tokens and Fast at $0.08; image tokens cost $0.40 per million for either. The model documentation and dated release note describe a shared space, 128K context, text/image inputs and more than 100 languages. The launch announcement supplies the prices; these are vendor rates, not a measured workload bill.
| Property | Pro | Fast |
|---|---|---|
| Text / 1M tokens | $0.12 | $0.08 |
| Image / 1M tokens | $0.40 | $0.40 |
| Context | 128K tokens | 128K tokens |
| Vendor positioning | Quality-focused indexing and retrieval. | Latency-sensitive query traffic. |
| Vector dimensions | 256–2048, six supported sizes. | Same supported sizes. |
The supported dimensions are 256, 512, 768, 1024, 1536 and 2048, with float, int8 and binary outputs. Select a consistent representation for the index and query path. A vector database may impose its own requirements on dimensions, distance metric and supported formats.
02 — Practical implicationsWhat a shared space changes
An embedding is a numerical representation used to compare a query with stored material. When two models produce compatible representations, the index built with one can be queried with the other as documented. Cohere recommends Pro for indexing and Fast for queries, which separates an infrequent corpus operation from repeated interactive traffic.
Prepare the searchable material
Test Pro on the document collection, including difficult layouts and languages.
Represent the user’s question
Compare Fast and Pro against the same index and relevance judgments.
Use the retrieved evidence
Check that the final answer cites and uses the correct retrieved material.
Do not generalize compatibility beyond the stated family. An existing index made with another embedding model should be treated as a separate migration. Matching vector length alone does not prove that distances have the same meaning. Keep the old index until the replacement has passed retrieval and application checks.
03 — Practical implicationsLonger inputs do not eliminate document design
A 128K context window allows larger inputs, but the best retrieval unit may still be a section, page or other coherent chunk. A whole document can contain several unrelated answers. If the stored representation blends them together, the system may retrieve a broadly relevant document without locating the passage needed for a precise response.
Compare chunking strategies on actual questions. Include queries that require a table row, an exception clause or information near the end of a long document. Preserve identifiers that let the application link a retrieved vector back to the right source location. A high similarity score is less useful when the user cannot inspect the evidence.
For visually structured documents, compare parsed text with the supported image or mixed-input path. Use the same relevance judgments and keep parsing failures visible. The model cannot recover information that a preprocessing step removed before embedding.
04 — Practical implicationsRun a small parallel-index evaluation
| Step | Evidence to keep |
|---|---|
| Freeze inputs | A versioned document sample and representative query set. |
| Build a candidate | Model, dimensions, format, chunking and source identifiers. |
| Judge retrieval | Relevant results and important misses for each query. |
| Measure serving | Latency, throughput and actual usage under the expected load. |
| Test rollback | A verified route back to the previous index and query model. |
Use human relevance judgments where the answer matters, and inspect failures by category. An average can hide poor results for one language or one document format. Keep permissions and metadata filters in the evaluation too: a relevant document that the user is not allowed to access is a failed application result.
Illustratively, embedding 100 million text tokens at the listed rates costs $12 with Pro or $8 with Fast, excluding every other charge. That $4 difference may be small beside the cost of parsing, storing and reviewing the corpus. For frequent query traffic, latency and repeated usage can matter more. Recalculate from your actual volumes.
05 — Practical implicationsChoose the pair that improves useful retrieval
The model-routing guide explains accepted-result economics; the agent-memory guide places retrieval in a larger system. Keep the access boundary fixed while testing. Our AI transformation service supports those controlled evaluations.
Compare query quality before replacing the existing index
Use the shared space to test a practical Pro/Fast pairing, then choose from relevance, latency and complete operating cost. Preserve the old index until the new pipeline handles both useful answers and difficult misses correctly.