Model comparisons · 2026-10-01
What is Voyage Rerank 3? API pricing, cost calculation and how it differs from Rerank 2.5
Use Voyage Rerank 3 and Rerank 3 Lite through LLMTR: how they differ from Rerank 2.5, a cost calculation by query and document tokens, the limits and a first /v1/rerank request.
Short answer: what is Voyage Rerank 3?
Voyage Rerank 3 is Voyage AI's reranking model: it reads a query and a list of candidate documents together and produces a relevance score for each document. Voyage positions it as the most accurate member of its reranker family and the default choice for most applications. The same generation also has a Lite version that is faster and cheaper for latency-sensitive work. On LLMTR both models are called as voyageai/rerank-3 and voyageai/rerank-3-lite with your existing LLMTR API key; no separate Voyage account is needed.
A reranker does not replace embedding search, it runs behind it. Embedding search vectorizes the query and the document separately, so it is fast across millions of records but can miss subtle differences in meaning. A reranker reads the query and each candidate at the same time. That is why it orders more accurately, but because it evaluates every candidate separately it is usually applied only to the top 50-100 candidates.
How Rerank 3, Rerank 3 Lite and Rerank 2.5 differ
On price the new generation sits at the same tier as the old one: Rerank 3 costs the same as Rerank 2.5, and Rerank 3 Lite the same as Rerank 2.5 Lite. The query ceiling and the per-request document limit did not change either. Voyage says the new models are better than the old ones in every respect, such as quality, latency and throughput, and now lists the 2.5 models among its older models. LLMTR did not measure this comparison itself.
The only functional difference visible in the documentation is instruction support. Voyage documents steering the relevance criterion by adding an instruction to the query only for the 2.5 models; there is no such statement for Rerank 3. All four models are available on LLMTR.
| Field | Rerank 3 | Rerank 3 Lite | Rerank 2.5 | Rerank 2.5 Lite |
|---|---|---|---|---|
| Price (1M tokens) | $0.05 | $0.02 | $0.05 | $0.02 |
| Context (query + one document) | 32,000 tokens | 32,000 tokens | 32,768 tokens | 32,768 tokens |
| Query ceiling | 8,000 tokens | 8,000 tokens | 8,000 tokens | 8,000 tokens |
| Documents per request | 1,000 | 1,000 | 1,000 | 1,000 |
| Instructions in the query (Voyage docs) | Not documented | Not documented | Yes | Yes |
Cost calculation: the query is counted again for every document
Voyage reranking bills on total processed tokens, and the total is computed as (query tokens × document count) + the sum of tokens across all documents. In other words, the query is counted once more for every document you rerank. The usage.total_tokens field in the response carries this total, and the charge is computed from that number. Reranking has no output tokens and no cache discount.
As an example, take a 20-token query reranking 50 candidates of 300 tokens each. The total is 20 × 50 + 50 × 300 = 16,000 tokens. That request costs 16,000 × 0.05 / 1,000,000 = $0.0008 on Rerank 3 and $0.00032 on Rerank 3 Lite. Over a thousand requests the difference is $0.80 against $0.32. Raising the candidate count from 50 to 100 doubles the cost; choose how many candidates to rerank by the accuracy you gain.
No margin is added to model prices; LLMTR's platform fee is taken only when you top up credit.
First request: /v1/rerank
The request goes to LLMTR's /v1/rerank endpoint with your LLMTR API key. The rerank endpoint is not part of the OpenAI SDK; call it over plain HTTP. top_k returns only the N highest-scoring results; top_n is also accepted for clients written against the Cohere spelling. When return_documents is true, each result also carries its own text.
The data field in the response comes back in descending order of relevance score. index is the document's position in the list you sent; use it to match results to your own records. Scores are not comparable across models, so do not carry your old score threshold over unchanged when moving from Rerank 2.5.
A Voyage Rerank 3 request for POST /v1/rerank
{
"model": "voyageai/rerank-3",
"query": "When are invoices issued?",
"documents": [
"Invoices are issued on the first business day of each month.",
"Shipping takes 2-3 business days.",
"The return period is 14 days."
],
"top_k": 2
}
Moving from Rerank 2.5
Voyage has not shut down the 2.5 models and they stay available on LLMTR. Moving is not required, but starting a new integration on Rerank 3 makes sense. When you move, check these four points:
- Change the model field from voyageai/rerank-2.5 to voyageai/rerank-3, and from voyageai/rerank-2.5-lite to voyageai/rerank-3-lite; the request and response shapes are the same.
- If you cut results at a score threshold, set the threshold again on the new model with a few real queries.
- If you add instructions to the query, run the same queries on both models and compare the ordering; Voyage documents that usage only for 2.5.
- For long inputs, truncation is on by default and cuts text that exceeds the limit; if you send false, a request over the limit returns an error.
Which model to pick, and when
Start with Rerank 3 in RAG flows where ranking accuracy directly decides answer quality, for example legal, financial or technical document search. Rerank 3 Lite can be enough for search boxes where every user query is reranked, in-chat retrieval, and work where volume drives cost.
Before deciding, build a small evaluation set from your own queries, run both models on the same candidates, and compare the accuracy of the top five results with the total cost. For Turkish-heavy content, qwen/qwen3-reranker-8b in the catalog is called through the same endpoint and can join the same comparison.
Frequently asked questions
How much does the Voyage Rerank 3 API cost?
Rerank 3 costs $0.05 and Rerank 3 Lite $0.02 per 1M tokens. The charge is computed on (query tokens × document count) + the sum of document tokens; there are no output tokens.
Can I move my Rerank 2.5 integration to Rerank 3?
Yes. Changing the model field is enough; the request and response shapes are the same. If you use a score threshold, tune it again on the new model. The Rerank 2.5 models stay available on LLMTR.
How many documents can Rerank 3 rerank in one request?
At most 1,000 documents. The query can be at most 8,000 tokens, the query plus any single document at most 32,000 tokens, and the request total cannot exceed 600,000 tokens. For more, split the list into batches and merge the returned scores on your side.