Retrieval quality depends on the embedding model, and cost depends on knowing what each call costs. This release adds a provider on one side and a price on the other.
Cohere
- Embeddings The four third-generation embedding models on the second-generation API, with retry, parallelism and a configurable input type.
- Tokenizer Encode, decode and estimate through the tokenize and detokenize endpoints, exposed to scripts, with credentials as a plain key, an encrypted field or a third-party credential.
Vertex and pricing
- Multimodal embeddings Structured requests for text, image, audio, video and documents with task-specific parameters and usage tracking, on the global endpoint, with the deprecated models replaced.
- A price per model Request and response pricing fields per provider model, and a service that answers what a model costs. Builder factories for chat, embedding, rerank and speech models are now global.
The spend quota of December 2024 needed a price to count. This is where the price comes from.