Model comparisons · 2026-09-30
What is GPT-6.1 Sol? API pricing, cache cost and how it differs from GPT-6 Sol
Use OpenAI GPT-6.1 Sol through LLMTR: how it differs from GPT-6 Sol, a cost calculation with the $0.10 cache-read price, reasoning levels and a first Responses request.
Short answer: what is GPT-6.1 Sol?
GPT-6.1 Sol is the model OpenAI offers as the improved version of GPT-6 Sol. OpenAI positions it for agentic coding, computer use, and professional work such as complex documents and multi-step workflows. On LLMTR it is called as openai/gpt-6.1-sol with your existing LLMTR API key; no separate OpenAI account is needed.
The model has a 1,050,000-token context window and a 128,000-token output ceiling. It accepts text, image and PDF input; audio input is not supported. Function calling, structured output with strict json_schema, and streaming are supported. Input and output are priced the same as GPT-6 Sol; the difference is in cached input.
How does it differ from GPT-6 Sol?
The only difference in the price table is the cache read: on GPT-6.1 Sol a token read from the cache bills at one twentieth of the input price, on GPT-6 Sol at one tenth. Cache write, input and output are the same on both models.
The other difference is on the reasoning side: GPT-6.1 Sol does not accept the none level, and its lowest level is low. If you have a latency-sensitive flow that needs reasoning switched off entirely, GPT-6 Sol remains the right choice for that work. Both models are available on LLMTR.
| Field | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|
| Input | $2.00 | $2.00 |
| Cache read | $0.10 | $0.20 |
| Cache write | $2.50 | $2.50 |
| Output | $10.00 | $10.00 |
| Reasoning levels | low, medium, high, xhigh, max | none, low, medium, high, xhigh, max |
| Context | 1,050,000 tokens | 1,050,000 tokens |
The evaluation results OpenAI published
The results below come from OpenAI's GPT-6.1 Sol announcement; LLMTR did not rerun these tests. An evaluation score does not guarantee the result on your own workload: before deciding, run the same tasks on both models and compare accuracy and total cost.
- DeepSWE v1.1, software engineering in real codebases: the same score as GPT-6 Astra at roughly one fifth of the cost according to OpenAI, and 6.4 points above GPT-6 Sol's best score.
- AutomationBench, multi-step workflows: 4.8 points above GPT-6 Sol at the medium reasoning level.
- OSWorld 2.0, computer use: seven points above GPT-6 Sol at the highest reasoning level.
- Factual accuracy: at xhigh, the factual error rate drops from GPT-6 Sol's 4.5% to 4.1%.
Cost calculation: first request and repeated request
Writing to the cache is not free on the GPT-6 family. On the first request, the cacheable prefix of the prompt bills at 1.25x the input price, which is $2.50 per 1M tokens. On later requests with the same prefix, those tokens are read at $0.10. Cache-write and cache-read tokens are a subset of input tokens; the same token is never counted twice.
As an example, take an agent step with 10,000 input tokens: 8,000 are fixed system instructions and tool definitions, 2,000 change on every step, and the response produces 1,000 output tokens. The first request costs (2,000 × 2.00 + 8,000 × 2.50 + 1,000 × 10.00) / 1,000,000 = $0.034. Once the prefix is read from the cache, the same step costs $0.0148. The same repeated request on GPT-6 Sol costs $0.0156; the difference comes entirely from the cache read and grows with the prefix.
On requests whose input exceeds 272,000 tokens, the whole request bills at the long-context rates: input and cache rates double, and the output rate rises to 1.5x. Reasoning tokens bill at the output price. No margin is added to model prices.
First request: Responses API
The request goes to LLMTR's Responses endpoint with your LLMTR API key. The same model can also be called through Chat Completions, and you can send tool definitions together with a reasoning level on that endpoint too.
At high reasoning levels, do not keep max_output_tokens low: if the budget is used up on reasoning, the response text can come back empty. On your first try, read the usage field in the response; cache-write and cache-read counts appear there separately, and the charge is computed from those counts.
A GPT-6.1 Sol request for POST /v1/responses
{
"model": "openai/gpt-6.1-sol",
"reasoning_effort": "medium",
"input": "List the tests that the changes in this pull request could affect.",
"max_output_tokens": 2048
}
Which model to pick, and when
Start with GPT-6.1 Sol for writing and debugging code, tool-using agents, and document-heavy work. In agent flows that send the same context again and again, the cache-read price makes this model cheaper than GPT-6 Sol.
Use GPT-6 Sol if you need reasoning switched off entirely. For the hardest end-to-end work, look at GPT-6 Astra, which OpenAI positions as its most capable model; for high-volume, cost-sensitive work, look at GPT-6 Luna. The Ultrafast version OpenAI announced alongside it is not offered on LLMTR; if a request carries the service_tier field, it runs on the standard tier and bills at the standard price.
Frequently asked questions
How much does the GPT-6.1 Sol API cost?
Input is $2.00, cache read $0.10, cache write $2.50 and output $10.00 per 1M tokens. On requests whose input exceeds 272,000 tokens, input is $4.00, cache read $0.20, cache write $5.00 and output $15.00.
Can I move my GPT-6 Sol integration to GPT-6.1 Sol?
Yes. Change the model field from openai/gpt-6-sol to openai/gpt-6.1-sol. The one exception is the reasoning level: if you send none, the request returns 400; pick low or a higher level.
Does GPT-6.1 Sol support temperature?
No. temperature and top_p are not supported on this model. Instead of rejecting a request that contains them, LLMTR removes the fields and reports them in the response's llmtr_dropped_parameters field.