RAG and data · 2026-10-06

Voyage Models With Quantization and Flexible Dimensions

Compare seven flexible Voyage embedding models on LLMTR, choose 256 to 2048 dimensions, select float, int8, uint8 or binary, and size a RAG index.

Simple diagram of document and query inputs becoming Voyage embedding vectors, with arrows comparing flexible dimensions and quantized vector storage.

Short answer: Seven flexible Voyage models

LLMTR (llmtr.com) is a gateway that gives access to language models hosted in Türkiye and at global providers through one OpenAI-compatible API. Use https://llmtr.com/v1 as the OpenAI-compatible base URL and POST embeddings to https://llmtr.com/v1/embeddings; voyageai/voyage-4 supports four output sizes and multiple quantization formats for vector search.

The flexible models are voyageai/voyage-4-large, voyageai/voyage-4, voyageai/voyage-4-lite, voyageai/voyage-code-4, voyageai/voyage-code-3, voyageai/voyage-context-4 and voyageai/voyage-multimodal-3.5. The two remaining models, voyageai/voyage-finance-2 and voyageai/voyage-law-2, return fixed 1024-dimension float vectors instead. Their context is 32,000 tokens, except voyageai/voyage-law-2 at 16,000; the table also records output sizes, quantization support and input-token prices.

Voyage embedding models on LLMTR (2026-10-06)
Model idUSD per 1M input tokensContextOutput dimensionsQuantization
voyageai/voyage-4-large$0.1232,000256, 512, 1024, 2048Yes
voyageai/voyage-4$0.0632,000256, 512, 1024, 2048Yes
voyageai/voyage-4-lite$0.0232,000256, 512, 1024, 2048Yes
voyageai/voyage-code-4$0.1232,000256, 512, 1024, 2048Yes
voyageai/voyage-code-3$0.1832,000256, 512, 1024, 2048Yes
voyageai/voyage-context-4$0.1832,000256, 512, 1024, 2048Yes
voyageai/voyage-multimodal-3.5$0.1232,000256, 512, 1024, 2048Yes
voyageai/voyage-finance-2$0.1232,0001024No (float only)
voyageai/voyage-law-2$0.1216,0001024No (float only)

What Voyage models offer quantization and flexible dimensions?

On LLMTR, seven Voyage embedding models support 256, 512, 1024 and 2048 output dimensions plus quantization: voyageai/voyage-4-large, voyageai/voyage-4, voyageai/voyage-4-lite, voyageai/voyage-code-4, voyageai/voyage-code-3, voyageai/voyage-context-4 and voyageai/voyage-multimodal-3.5. Each listed model accepts the same four output sizes and supports quantized output formats through the LLMTR embeddings API.

These models accept float, int8, uint8, binary and ubinary through output_dtype. voyageai/voyage-finance-2 and voyageai/voyage-law-2 are different: they return fixed 1024-dimension float vectors and do not support quantized output. Choose them when the integration requires a 1024-dimension float vector; dimension and dtype requests outside those values are refused.

Dimensions, dtypes and request validation

In the request below, dimensions sets the output vector width to 512, while output_dtype selects int8. LLMTR also accepts output_dimension and output_dimensionality as aliases for dimensions; the endpoint is POST https://llmtr.com/v1/embeddings. The model field selects voyageai/voyage-4 in this exact request. Both dimension field names select the same output size.

Set input_type to query when searching and document when indexing. Flexible models accept 256, 512, 1024 or 2048 for dimensions; unsupported dimension or output_dtype values return 400 invalid_request, with a message listing the accepted values. For voyageai/voyage-finance-2 and voyageai/voyage-law-2, dimensions other than 1024 and output_dtype values other than float are also refused with 400.

This request uses voyageai/voyage-4 to create a 512-dimensional int8 document embedding through the OpenAI-compatible embeddings endpoint.

curl https://llmtr.com/v1/embeddings \
  -H "Authorization: Bearer $LLMTR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "voyageai/voyage-4",
    "input": ["Invoices are due 30 days after the delivery date."],
    "input_type": "document",
    "dimensions": 512,
    "output_dtype": "int8"
  }'

What quantization saves for 1M vectors

The table estimates storage for 1,000,000 embeddings at each supported flexible dimension. At 1024 dimensions, float requires about 4.1 GB at 4 bytes per value, int8 requires 1.0 GB at 1 byte and binary requires 128 MB at 1 bit; at 512 dimensions, int8 requires 512 MB.

Quantization changes index size and search speed, not catalog price. Billing uses input tokens only: 10M tokens cost $0.60 on voyage-4 and $0.20 on voyage-4-lite, whether the output uses float, int8 or binary. Selecting a smaller dimension or a supported quantized dtype does not change the charge.

Index size for 1,000,000 vectors
Dimensionsfloat (4 bytes)int8 (1 byte)binary (1 bit)
20488.2 GB2.0 GB256 MB
10244.1 GB1.0 GB128 MB
5122.0 GB512 MB64 MB
2561.0 GB256 MB32 MB

How to choose dimensions and output dtypes

For most RAG integrations, start at 1024 dimensions with int8. Test 512 dimensions or binary embeddings on your own query set before switching because recall can drop. Binary embeddings are often used for a fast first pass followed by rescoring with full vectors.

Embed queries and documents with the same model, dimensions and output_dtype at index time and at query time; vectors of different sizes cannot be compared. The seven flexible models can use 256, 512, 1024 or 2048 dimensions. By contrast, voyageai/voyage-finance-2 and voyageai/voyage-law-2 always return 1024-dimension float vectors.

Getting started on LLMTR: account, credit, key and first request

Step 1: Create an account at llmtr.com and verify your email. Step 2: Add credit from the Dashboard through the secure checkout. LLMTR adds an 8% platform margin once, on top of the requested top-up amount, so a $10.00 credit is charged $10.80. The platform margin is not added to model prices; every API call is deducted at the model's catalog price.

Step 3: Create an API key under Dashboard > API Keys. The raw key is shown only once, so store it in an environment variable such as LLMTR_API_KEY. Step 4: Send the first request to https://llmtr.com/v1 with any OpenAI-compatible SDK by changing only base_url and api_key. The Dashboard usage page shows usage and spend per key. The same key and credit balance work for every catalog model; switching models means changing the model string.

Create access and call Voyage embeddings

Create an LLMTR account, add credit, create an API key and call the Voyage embeddings endpoint with an OpenAI-compatible SDK.

  1. Account. Create an account at llmtr.com and verify your email.
  2. Credit top-up. Add credit from the Dashboard through the secure checkout. An 8% platform margin is added once, so a $10.00 credit is charged $10.80; model prices do not include the margin.
  3. API key. Create an API key under Dashboard > API Keys. Store the raw key in an environment variable such as LLMTR_API_KEY because it is shown only once.
  4. First request. Send an embedding request to https://llmtr.com/v1/embeddings using any OpenAI-compatible SDK by changing only base_url and api_key.

Frequently asked questions

Which Voyage models support flexible dimensions and quantization?

The seven are voyageai/voyage-4-large, voyageai/voyage-4, voyageai/voyage-4-lite, voyageai/voyage-code-4, voyageai/voyage-code-3, voyageai/voyage-context-4 and voyageai/voyage-multimodal-3.5. Each accepts 256, 512, 1024 and 2048 dimensions and quantized output formats.

Does int8 or binary embedding change the LLMTR price?

No. LLMTR bills embeddings by input tokens only, so int8 and binary do not change the price; 10M tokens cost $0.60 on voyage-4 and $0.20 on voyage-4-lite.

Can I use 768 dimensions with Voyage embeddings on LLMTR?

No. Flexible Voyage models accept 256, 512, 1024 or 2048; 768 returns 400 invalid_request. The fixed voyageai/voyage-finance-2 and voyageai/voyage-law-2 models require 1024 dimensions and float output.

Related posts