Qwen / qwen/qwen3-reranker-8b

Qwen3 Reranker 8B - access through LLMTR

Qwen3 Reranker 8B evaluates a query together with a list of candidate documents and produces a relevance score between 0 and 1 for each one. Embedding search vectorizes query and document separately for speed; a reranker reads both at once and therefore orders results more accurately. The typical pattern is to retrieve the top 50-100 candidates with an embedding model, rerank them, and pass only the most relevant few to the language model, which improves answer quality and lowers the generation model's token cost. It reads over 100 languages, including Turkish, and accepts up to 32,768 tokens per query-document pair. It is called through /v1/rerank and produces no text. Billing is on total processed tokens: query and document tokens are counted together for every pair.

Technical specifications

Canonical IDqwen/qwen3-reranker-8b
ProviderQwen
Context window32,768 tokens
OperationsRERANK
Modalitiestext

Pricing

An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.

OperationMetricUnitPrice
RERANKINPUT_TEXTPER_1M_TOKENS$0.050000

Example usage

With existing OpenAI SDK flows, change only the base URL and model identifier.

curl https://llmtr.com/v1/chat/completions \
  -H "Authorization: Bearer llmtr-your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen/qwen3-reranker-8b","messages":[{"role":"user","content":"Hello"}]}'

Related models