Cohere launches Embed 5 embedding models

Pro and Fast tiers go generally available Sept. 30 with a shared vector space, 128K context and multimodal inputs

Cohere on Wednesday released Embed 5, a two-tier family of embedding models the company said is built for enterprise retrieval. The models are embed-v5.0-pro and embed-v5.0-fast.

The blog post is blunt about the pitch. “Today, we’re releasing Embed 5, a new family of embeddings models at the frontier of high-quality enterprise retrieval.” The product changelog calls the pair “Cohere’s most powerful embeddings family yet.”

Both IDs share one embedding space. Index with one. Query with the other. Cohere’s recommended split is Pro for the corpus and Fast for the live query path. Context length is 128,000 tokens on each. Inputs are text, images, and mixed text-and-image pages, including PDFs. Language coverage is listed at more than 100.

Output width is flexible. Dimensions: 2,048, 1,536, 1,024, 768, 512, 256. Default on the model card is 2,048. Formats: float, int8, binary. Matryoshka representations are on. Similarity metrics listed in the docs are cosine, dot product and Euclidean distance.

Pricing on the announcement page is $0.12 per million text tokens for Pro and $0.08 per million for Fast. Image tokens are $0.40 per million on both. The models are generally available on the Cohere API and Model Vault, Microsoft Foundry and Amazon SageMaker. Private serving is listed through Model Vault and vLLM.

Cohere published its own scorecard. On ViDoRe V3, it put Embed 5 Pro at an 85.8 average, an 8.8-point gain over Embed 4. Fast was listed at 84.5. The same table placed Voyage 4 Large at 83.7, Gemini Embedding 2 at 83.2 and OpenAI text-embedding-3-large at 75.5. Those figures are Cohere’s tests, not a third-party audit.

Finance numbers sit in the same post. FinanceBench: Pro 80.1, Fast 80.0. FinQA: Pro 90.0, Fast 88.8. ViDoRe V3 Finance: Pro 85.0, Fast 83.9. On a parsed-PDF suite, Cohere listed Pro at 84.8 and Fast at 83.4, ahead of its own Embed 4 score of 78.6. Fused text-image retrieval: Pro 82.3 average across five datasets, Fast 81.2. Page-image retrieval on financial pages: Pro 77.0, Fast 73.2.

Multilingual averages for German, French, Spanish, Italian and Russian were listed at 77 for Pro, about seven points above Embed 4. The company also said Embed 5 is the first model family it evaluated with RCP-nDCG@10, a retrieval metric it published the same day.

Storage math is in the blog, not the changelog. “A 2,048-dimensional float32 vector requires 8 KB; a 1,024-dimensional int8 vector uses 1 KB; and a 256-dimensional binary vector just 32 bytes—a 256x reduction.” Across 100 million chunks, Cohere said that cut takes raw vector storage from about 819 GB to 3.2 GB. “For most deployments, we recommend 1,024-dimensional int8 vectors as the ideal performance-efficiency point.”

The company pointed the models at search, retrieval-augmented generation and agent loops. Named verticals on the post include financial filings, retail catalogs, legal archives and customer-service search. Batch embedding is documented for large ingest jobs. Integrations listed include LangChain, Haystack, Weaviate, Qdrant, Pinecone, Elasticsearch, MongoDB, Redis, Milvus and OpenSearch.

Cohere said it will discuss Embed 5 and Parse 5 on X on Oct. 8. Compass Cloud remains a separate managed-search beta on the company’s site.

No executive is quoted by name in the launch post. The model IDs, the prices and the score tables are what the company put on the page.

Subscribe — you own it

No tracking, no middleman. Follow by RSS (nothing is collected) — or add your email to our self-hosted list.

RSS feed →
Subscribe
Notify of
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x