Model comparison · 2026-08-28
Upstage Solar Pro 4 versus Pro 3: context, reasoning and price
Compare Solar Pro 4 and Solar Pro 3 by context, separate output limits, explicit reasoning settings and dated token prices using one fair evaluation task.
Should you use Solar Pro 4 or Pro 3?
Solar Pro 4 provides more room for large document sets or a clearly defined long-output ceiling. If the smaller context is enough, do not infer quality from the version number. Evaluate Pro 3 on the same task and rubric. No new live call or superiority result appears here.
The LLMTR identifiers are upstage/solar-pro4 and upstage/solar-pro3. Both use text-based Chat Completions. Extract text from PDFs or images first and preserve source locations. Compare them with fictional rather than personal or confidential data.
Read context and output limits separately
Upstage cards use rounded 512K and 128K labels. LLMTR's August 11, 2026 provider-rejection measurements give exact limits: Pro 4 has 524,288 context and a separate 131,072 maximum output; Pro 3 has 131,072 context and no separate maximum-output field.
Do not turn Pro 3's context into a response ceiling. Prompt, reasoning and final answer affect the budget. Give both models the same max_tokens value and reject runs ending with finish_reason length as incomplete.
| Criterion | Solar Pro 4 | Solar Pro 3 |
|---|---|---|
| LLMTR identifier | upstage/solar-pro4 | upstage/solar-pro3 |
| Context | 524,288 tokens | 131,072 tokens |
| Separate maximum output | 131,072 tokens | Not published; do not derive it from context |
| Input scope | Text | Text |
Make reasoning settings explicit
LLMTR accepts none, minimal, low, medium, high, xhigh and max for Pro 4; Pro 3 accepts none, minimal, low, medium and high. Do not send xhigh or max to Pro 3. Recorded reasoning begins at low on Pro 4 and medium on Pro 3.
Current Upstage documentation says omission enables reasoning on Pro 4 and disables it on Pro 3. The Pro 4 statement conflicts with LLMTR's August 11 measurement. Avoid defaults: send minimal to both for no reasoning and medium to both for reasoning.
| Condition | Solar Pro 4 | Solar Pro 3 | Purpose |
|---|---|---|---|
| A | minimal | minimal | Compare without reasoning tokens |
| B | medium | medium | Compare with reasoning enabled |
| Unused Pro 4 levels | none, low, high, xhigh, max | — | Outside this evaluation |
| Unavailable on Pro 3 | — | xhigh, max | Do not send |
Calculate price within its effective window
Raw rates are USD per million tokens. Pro 4's promotion ends September 10, 2026 at 23:59 UTC; standard rates start September 11 at 00:00 UTC. Pro 3 rates were verified August 28. No credit top-up margin is included.
Calculate uncached input, cache reads and total output separately. Reasoning is output; do not add reasoning_tokens again when completion_tokens gives the total. Price ordering changes after the promotion, so do not base a long-term choice on today's discount.
| Model and period | Input | Cache | Output |
|---|---|---|---|
| Pro 4, before 2026-09-11T00:00Z | 0.03 | 0.006 | 0.12 |
| Pro 4, at and after 2026-09-11T00:00Z | 0.30 | 0.06 | 1.20 |
| Pro 3, verified August 28, 2026 | 0.15 | 0.015 | 0.60 |
Use the same document-reconciliation task
Find amount, date and delivery-term differences between two fictional contract summaries. Give both models the same text, instruction, order and 4,096 max_tokens budget. Run minimal then medium without repairing the prompt. The text below is the shared prompt, not a model response.
Fictional task sent unchanged to both models
Use only the two sources below. For every difference, state the field, Source A value, Source B value and supporting sentence. Do not infer information absent from a source.
Source A: The total is TRY 240,000. Delivery is October 15, 2026. The delay penalty is 0.2 percent per day.
Source B: The total is TRY 245,000. Delivery is October 22, 2026. No delay penalty is stated.
Finish by answering only these three checks: Were all numeric differences found? Was the missing term clearly marked? Is there any claim outside the sources?
Freeze the rubric before generating results
Hide the model name and score each answer 0, 1 or 2 on five criteria. Keep conditions separate and do not select only successful examples. Use equal repeat counts and alternate model order. No score, speed or cost result is reported here.
- Coverage: are both the amount and delivery-date differences present?
- Missing information: is the absent penalty term in Source B explicit?
- Evidence: is every value tied to the correct source sentence?
- Fidelity: did the answer add an amount, date or term absent from the sources?
- Completion: are all requested fields present and is the response untruncated?
Decide from task fit, budget and failure type
A task outside Pro 3's context is unsuitable for Pro 3; within it, Pro 4's newer number is not selection evidence. Record score, input, cache reads, output, reasoning tokens and finish reason. Do not log prompts or answers in production.
Set the acceptance threshold first. If both models pass, compare cost using the real token mix and effective price period. A lower rate does not make a failed task usable.
Frequently asked questions
Is Solar Pro 4 better than Pro 3 on every task?
No. General evaluations do not guarantee your workload; measure both with the same prompt and a rubric fixed in advance.
Can Solar Pro 3 output up to 131,072 tokens?
Do not treat it as a separate output ceiling. It is the context value; Pro 3 publishes no separate maxOutputTokens value.
Can I compare the models with default reasoning?
No. Sources disagree on the Pro 4 default. Explicitly send minimal and medium conditions to both models.