Anthropic released Claude Sonnet 5.5 on September 28, 2026, and the headline numbers look extraordinary: 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%, and even beating Opus 5.5's 66.4% at half the per-token price. But the 30% cost savings Anthropic advertises only applies at default effort settings. Independent testing from Artificial Analysis found that at maximum effort, Sonnet 5.5 actually costs $7.60 per task — roughly 49% more than Sonnet 5's $5.09, not 30% less. The benchmark gains are real. The "cheaper" claim depends entirely on how you use it.
What actually shipped
Anthropic launched Claude Sonnet 5.5 on September 28 as the second model in its Claude 5.5 family, following Opus 5.5's September 22 debut. The company positioned it as "a faster, lower-cost complement to Claude Opus 5.5," optimized for "well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets".
Here are the benchmark results Anthropic published:
Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | — |
GDPval-AA v2.1 (knowledge work) | 1,844 | 1,449 | 1,846 | 1,487 |
AA-Briefcase v1.1 | 1,811 | 1,359 | 1,822 | 1,483 |
CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
OSWorld 2.1 (computer use) | 80.1% | 57% | 81.8% | — |
Source: Anthropic's official Sonnet 5.5 announcement.
On Terminal-Bench 4.0 — a test of whether an AI agent can complete complex professional tasks by typing commands on its own — Sonnet 5.5's 70.6% would make it the highest-scoring model Anthropic has ever released on that specific test, beating even the more expensive Opus 5.5. On knowledge-work evaluations like GDPval-AA, which covers real-world tasks across 44 occupations, Sonnet 5.5 reaches near-parity with Opus 5.5, trailing by only two points.
Anthropic's own framing, however, is more cautious than the benchmark table suggests: "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment".
The cost paradox: "30% cheaper" is conditional
Anthropic's announcement leads with cost savings: Sonnet 5.5 "runs 30%+ faster, and costs up to 30% less for most work". The API pricing is unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output tokens. The company says the savings come from using fewer tokens and fewer tool calls to accomplish the same task, not from a lower per-token rate.
Independent testing tells a different story at the high end.
Artificial Analysis measured Sonnet 5.5 at maximum effort using approximately 193,000 output tokens per Intelligence Index task — the highest token usage they have ever recorded for any model, about 60% higher than Opus 5.5 (max) or Sonnet 5 (max), and roughly seven times GPT-6 Astra (max). The result: Sonnet 5.5 costs $7.60 per task on the Intelligence Index at max effort, compared to Sonnet 5's $5.09 — a 49% increase, not a 30% decrease.
The two claims are not contradictory. They are measuring different things. Anthropic's "30% cheaper" applies to well-scoped tasks at default or medium effort, where the model needs fewer tool calls and generates fewer tokens. Artificial Analysis measures the cost of achieving peak benchmark scores, which requires the model to "think" far longer — and Sonnet 5.5's adaptive reasoning at max effort burns through tokens aggressively.
Independent developer Simon Willison experienced the same pattern. He had Sonnet 5.5 build a small WebGL app at max effort, which consumed 128,000 tokens and cost $1.28. Switching to the xhigh effort tier, the same task finished in 41 seconds for about $0.06 — a 95% cost reduction.
The practical takeaway: Sonnet 5.5 is cheaper for everyday work at default settings. It is more expensive for complex tasks where you push it to its reasoning ceiling. Most migration defaults will land users in the first category. Power users who routinely run max effort will land in the second.
Better reasoning, or more thinking?

Terminal-Bench 4.0 tests agentic coding — navigating a terminal, running commands, debugging, completing multi-step software tasks. Sonnet 5.5's jump from 10.3% to 70.6% is not a marginal gain. It is a step change.
But the mechanism behind that gain matters. Artificial Analysis notes that Sonnet 5.5 achieves near-Opus performance on Terminal-Bench and GDPval by generating vastly more output tokens per task. The model is not necessarily "smarter" than its predecessor — it is working harder, for longer, to produce better answers on tests that reward thoroughness.
Where Sonnet 5.5 still lags Opus 5.5 is instructive. On AA-Omniscience, it scores 54% on factual accuracy against Opus 5.5's 66%, and it sits approximately six points lower on Humanity's Last Exam and SciCode compared to the flagship. These are the categories where raw capability — not effort — determines the outcome.
Anthropic's launch also carries a notable safety framing. Sonnet 5.5 is the first Sonnet model to ship with "frontier-style cyber safeguards" previously reserved for the company's most capable models, meaning higher-risk cybersecurity requests will visibly fall back to Sonnet 5. The company also added classifiers that block reasoning extraction, a response to an August 2026 incident in which researchers decoded 315,320 thinking blocks from public agent traces across OpenAI, Anthropic, and Google systems.
Anthropic says Sonnet 5.5 "does not advance the frontier of its models' capabilities" while simultaneously saying its cyber capabilities are "a large improvement warranting frontier-style safeguards". Both statements appear in the same announcement. A model can be stronger in one domain without being more capable overall — but the framing is selective by design.
What users are reporting
Early customer feedback has centered on speed. Box reported the new model is 2.4 times faster than the previous version, and Slack said it got better results without changing any prompts. One early tester described it as "fast at coding and can be steered quickly in iterative workflows".
But community discussion tells a more mixed story. On Chinese developer forums, several users reported that Sonnet 5.5's (thunderous thinking) at high effort settings makes it slower and more expensive than Opus 5.5 for complex tasks. One user noted: "The price of Sonnet 5.5 is half of Opus 5.5, but when actual work gets slightly more complex and triggers heavy thinking, it costs more than Opus 5.5 while doing worse".
Another thread regular reported that Sonnet 5.5 "doesn't consume that many tokens" for everyday use, and that their Pro-plan quota lasted all day "for the first time in months — no more mid-session rate-limit walls".
The pattern is consistent: Sonnet 5.5 is a clear win for well-scoped, everyday work. For complex tasks requiring sustained reasoning, Opus 5.5 remains the more efficient choice — both in cost and quality.
Where Sonnet 5.5 fits in model selection

Choose Sonnet 5.5 if:
You need fast iteration on well-scoped coding tasks — bug fixes, understanding unfamiliar codebases, targeted changes.
Your work involves creating documents, slides, and spreadsheets, including following existing templates.
You want near-Opus performance on agentic coding at half the per-token cost.
You value speed and responsiveness over sustained reasoning depth.
Choose Opus 5.5 if:
Your work requires complex, open-ended reasoning with sustained judgment.
You need the highest factual accuracy and scientific reasoning.
You can absorb higher per-token costs for better results on hard problems.
You want a model that does not burn 193,000 tokens to answer a difficult question.
Consider GPT-6 Sol if:
You are price-sensitive and your tasks do not require Anthropic's coding strengths.
You want more granular control over reasoning effort settings.
You need a model optimized for cost-efficient professional work.
On the Artificial Analysis Intelligence Index, Sonnet 5.5 scores 56 at max effort versus GPT-6 Sol's 47. But at high effort, the two models are nearly equivalent on intelligence at the same cost per task. The headline scores diverge more than the practical value does.
FAQ
Is Claude Sonnet 5.5 better than Opus 5.5?
It depends on the task. Sonnet 5.5 scores higher on Terminal-Bench 4.0 (70.6% vs. 66.4% at Opus 5.5's highest effort setting) and reaches near-parity on knowledge-work benchmarks. But Opus 5.5 remains stronger on factual accuracy (66% vs. 54% on AA-Omniscience) and scientific reasoning (approximately 6 points higher on Humanity's Last Exam and SciCode). Anthropic itself says Opus is "clearly stronger" for complex open-ended work.
Does Sonnet 5.5 actually cost less than Sonnet 5?
For well-scoped tasks at default or medium effort, yes — Anthropic's testing shows up to 30% lower cost per task. But at maximum effort on complex tasks, Artificial Analysis found Sonnet 5.5 costs $7.60 per Intelligence Index task versus Sonnet 5's $5.09, a 49% increase. The answer depends entirely on what effort level you use and what kind of work you do.
Why does Sonnet 5.5 use so many more tokens?
At max effort, Sonnet 5.5 generates approximately 193,000 output tokens per Intelligence Index task — the highest Artificial Analysis has ever measured for any model, roughly 60% higher than Opus 5.5 (max) and 7 times GPT-6 Astra (max). This is how it achieves near-Opus performance on benchmarks: by working longer and harder, not just by being more capable.
Should I switch from Sonnet 5 to Sonnet 5.5?
If you use Sonnet 5 for coding, bug fixing, or document creation at default settings, Sonnet 5.5 is a clear upgrade — the Terminal-Bench improvement from 10.3% to 70.6% is a step change, and speed gains are real. If you routinely run max effort on complex tasks, test the cost impact on your own workflows first. Anthropic's migration guidance says effort levels have been recalibrated and users should "re-run your effort sweep rather than carrying over a setting".
How does Sonnet 5.5 compare to GPT-6 Sol?
At max effort, Sonnet 5.5 scores 56 on the Artificial Analysis Intelligence Index versus GPT-6 Sol's 47. But at high effort, the two models are nearly equivalent on intelligence at effectively the same cost per task. Sonnet 5.5 offers stronger agentic coding; GPT-6 Sol offers more granular control over reasoning effort settings. Both are priced identically at $2/$10 per million tokens.
What's next for the Claude 5.5 family?
Claude Haiku 5.5, "built for high-volume and cost-sensitive applications," will join the family "in the coming weeks". That release will likely bring the 5.5 generation's improvements to the lowest price tier.
Letters
No letters yet — be the first to write.