Cohere has released Embed 5, a new embedding model family. It targets enterprise search, RAG, and agentic retrieval. The model family ships in 2 tiers. Embed 5 Pro targets maximum retrieval quality. Embed 5 Fast targets latency and cost on the live query path. Both accept text, images, and fused text plus image inputs. Both cover 100+ languages and read up to 128K tokens. The key design choice: Pro and Fast share 1 embedding space. You can index with one and query with the other.
Is it deployable today? Yes, both tiers are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker. Private VPC or on-prem serving runs through vLLM.
What Cohere Shipped
The API model IDs are embed-v5.0-pro and embed-v5.0-fast, per Cohere’s model docs. Both output 2048, 1536, 1024, 768, 512, or 256 dimensions, with 2048 as default. Embeddings come back as float, int8, or binary. Pro costs $0.12 per 1M text tokens. Fast costs $0.08. Image inputs cost $0.40 per 1M tokens on both.
Embed 5 can embed a page image directly. It can also fuse an image with its metadata into a single vector. That is important for scanned pages, slide decks, schematics, and charts, where text extraction drops information.
Pro and Fast: One Index, Two Query Paths
Cohere tested every corpus and query pairing across 40 development datasets. Normalized to Pro plus Pro at 100, a Pro index queried with Fast scored 98.4. An all-Fast setup scored 96.6. Cohere’s recommended pattern is to index with Pro and query with Fast. One constraint: both sides must use the same output dimension.
The split targets agentic workloads. An agent may issue dozens of searches per task, and query latency compounds. Cohere team reports Fast processed 377.3 documents per second versus 159.7 for Pro.
Benchmarks
On ViDoRe V3, Embed 5 Pro averages 85.8, an 8.8-point gain over Embed 4. Fast averages 84.5. Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2, and OpenAI text-embedding-3-large scores 75.5. On Cohere’s parsed-PDF suite, Pro leads at 84.8 against Voyage 4 Large at 83.6.
Finance is the strongest showing. Pro ranks first on FinanceBench (80.1), FinQA (90.0), and ViDoRe V3 Finance (85.0). Fast ranks second on all 3.
Multilingual results are mixed. Pro leads the European-language average at 77. However, Gemini Embedding 2 beats Pro on 9 of 10 further languages in Cohere’s own results table. Those include Japanese, Arabic, Hindi, and Telugu.
One important thing to note. Most numbers use RCP-nDCG@10, a new Cohere metric. It reorders a fixed candidate set, so it measures reranking quality more than first-stage retrieval. Cohere published the evaluation code, but independent replication is still pending.
Storage Costs at Scale
Embed 5 uses Matryoshka representation learning plus lower-precision outputs. A 2048-dim float32 vector needs 8 KB. A 1024-dim int8 vector needs 1 KB. A 256-dim binary vector needs 32 bytes. Across 100M chunks, raw storage drops from about 819 GB to 3.2 GB. Cohere recommends 1024-dim int8 as the default, citing near-full-precision quality.
