Frontier AI / VIDEO COMPANION

Claude Opus 5.5 vs GPT-6 Sol: thinking levels, benchmarks, and API cost

My release-day read on the official scores, thinking controls, long-context pricing, and what a same-prompt website demonstration can and cannot tell us.

EVIDENCE TYPEVendor analysis & qualitative demonstration

Video Source review

Which model would I use?

I’d start with GPT-6 Sol for routine API work where cost matters and the prompt stays under 272K input tokens. Its standard API rates are half of Opus 5.5’s for fresh input and output: $2 versus $4 per million input tokens, and $10 versus $20 per million output tokens. That gives me room to iterate, but it does not promise half the bill for every task. The models can use different numbers of tokens, and agent workflows add retries, tool calls, and cache behavior.

I’d test Opus 5.5 for coding work where its result justifies the spend. On the official FrontierCode 1.1 Main charts, Opus scored 54.6% at medium effort and GPT-6 Sol scored 45.9%; both publishers estimated $0.80 per task at those points. That 8.7-point difference is useful direction, not a controlled head-to-head result. The publishers ran their own evaluations, and I have not established identical harnesses, safeguards, or uncertainty intervals across those runs.

For a long document or agent session, I would price the whole request before choosing. Above 272K input tokens, Sol’s standard API rates change for the entire request. Its fresh-input price becomes $4 per million, matching Opus, while its output remains lower at $15 versus $20. The usual “Sol costs half as much” shortcut breaks there.

This is a September 22, 2026 release-day comparison of Anthropic’s Opus 5.5 announcement, OpenAI’s Sol announcement, and the vendors’ API documentation. I did not run an independent benchmark between the two models. The website exercise in my video is a separate, qualitative demonstration.

The API bill changes with context

These are standard, global-processing API rates in USD per million tokens as published on September 22. They exclude tool fees and taxes. Cache-write terms differ between the providers, so the equal $0.20 short-context cache-read price does not make their caching systems interchangeable.

Standard API rates · USD per million tokens
Token categoryOpus 5.5Sol up to 272K inputSol above 272K input
Fresh input$4.00$2.00$4.00
Output, including thinking tokens$20.00$10.00$15.00
Cache read$0.20$0.20$0.40
Cache write$5.00 for 5 minutes; $8.00 for 1 hour$2.50$5.00

The Sol threshold is based on input exceeding 272K tokens, and the higher rates apply to the full request, not just the tokens past the threshold. Anthropic lists the standard Opus rates across its 1M-token context window. Check Anthropic’s pricing and OpenAI’s model pricing again when you make a buying decision, because these are release-day rates.

Here is what the rate card does to three simple requests. These are arithmetic examples, not measured runs; the two models would not necessarily produce the same token counts.

Example request costs · USD per request
RequestOpus 5.5GPT-6 Sol
100K fresh input + 10K output$0.60$0.30
100K cached input + 10K fresh input + 10K output$0.26$0.14
400K fresh input + 10K output$1.80$1.75

Those totals omit initial cache writes, tool charges, and any extra agent turns. If a workflow spends most of its time reading a stable cached prefix, you need to measure its actual cache hits before projecting savings.

Thinking effort is part of the test configuration

Both models default to medium effort, but an effort label is a control within each provider’s system. It is not a shared compute budget. Opus 5.5 supports low, medium, high, xhigh, and max through adaptive thinking. GPT-6 Sol also supports those levels and none in the API. Opus does not offer a disabled-thinking setting for this model. Both models’ output limits include internal thinking or reasoning tokens as well as the visible answer. The model references explain the controls: Opus 5.5 and GPT-6 Sol.

The official FrontierCode charts show why I would not simply set every task to max:

FrontierCode 1.1 Main · score / estimated USD per task
EffortOpus 5.5GPT-6 Sol
Low47.3% / $0.4037.3% / $0.45
Medium54.6% / $0.8045.9% / $0.80
High54.0% / $1.0947.7% / $1.08
Xhigh51.4% / $2.2548.4% / $1.37
Max54.4% / $6.1949.3% / $2.14

Opus’s medium point is slightly higher than its max point on this chart, while its estimated cost jumps from $0.80 to $6.19. Sol improves from medium to max by 3.4 points, with estimated cost rising from $0.80 to $2.14. These are published points from separate runs, not a guarantee for your codebase. I would try medium first, examine the failure cases, and spend more effort only where the extra quality pays for itself. See the Anthropic and OpenAI charts for the source data.

Automation and computer use need tighter comparisons

On AutomationBench, Anthropic reports 40.0% at $1.37 per task for Opus 5.5 at max. OpenAI reports Sol’s best shown point at 33.2% at $0.27 per task at xhigh; Sol’s max point is 32.0% at $0.34. The score and price tradeoff is clear within each published sweep. The cross-vendor gap is only directional because the published setups and fallback accounting are not fully reconciled. A failure counted after a safeguard intervention is different from a result that can fall back to another model.

I would be even more careful with the computer-use figures. Anthropic’s Opus 5.5 OSWorld 2.0 result is 81.8% partial, while OpenAI reports Sol at 64.4% offline partial at max. Those labels describe different setups. Putting them in a single ranked table would make the comparison look firmer than it is. A missing result in either company’s launch material also means not reported, never zero.

What the website demonstration adds

In the video, I give both models the same landing-page task and inspect one attempt at each setting. Opus at high looked usable to me, and I preferred its xhigh page among the Opus results I could view. The Opus max attempt ran into my usage limit before producing a completed page, so there is no Opus max design to rank. On the Sol side, xhigh was my favorite; max also looked usable, while the Ultra orchestration missed visible text. Ultra used multiple agents and is a different workflow from the API effort settings in the tables above.

Those are my visual and usability judgments from a quick demonstration. I did not grade the pages with a repeatable rubric or verify production readiness. The account plans and product orchestration also differed. Identical task instructions do not equalize tools, thinking budget, or runtime. I would choose a promising page to continue editing, then check its behavior, content, and accessibility before shipping it.

The release data lead me to a workload decision rather than a universal winner. Start with Sol when the standard API price and iteration budget matter. Try Opus when the coding result is valuable enough to justify its higher token rates. For long-context work, recalculate with the actual input length and cache pattern before committing to either model.

In the full Opus 5.5 versus GPT-6 Sol video, I go through the official numbers and the website demonstration. Tell me in the YouTube comments which result you’d keep working on, or what workload you want compared next. Subscribe to DeepWakeLabs for the next test.

Claude Opus 5.5 and GPT-6-Sol are hosted models, so there are no official downloadable Hugging Face weights to link for this comparison. This article discusses vendor releases rather than local benchmark runs.

BRING YOUR SETUP TO THE CONVERSATION

What are you running?

Watch the comparison, then share your hardware and workload in the YouTube comments.

Watch & discuss on YouTube (opens in a new tab)
Back to the blog

DEEPWAKELABS