Claude Sonnet 5.5: 30% Faster, Same Price, 70.6% on Terminal-Bench
Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the 5.5 family and the fast, cheap sibling to Opus 5.5. The pitch is not raw power: it is cost per task. List pricing did not change — $2 per million input tokens and $10 output — but the company says the model generates text more than 30% faster and, because it needs far fewer tokens for the same work, comes out up to 30% cheaper per task. On benchmarks, it scores 70.6% on Terminal-Bench 4.0. Here is what that means in practice.
Quick answer: should you switch?
If you run Sonnet 5 in production, yes — same price table, faster output, fewer tokens burned. If you use Opus for simple tasks, move the routine ones to Sonnet 5.5 and keep Opus for heavy reasoning. If you use another vendor’s models, the math depends on your volume: at $2/$10, Sonnet 5.5 sits between Gemini 3.8 Flash (cheaper) and GPT-6 Astra (pricier and more capable).
Pricing and limits
| Item | Claude Sonnet 5.5 |
|---|---|
| Input | $2 per million tokens |
| Output | $10 per million tokens |
| Cache read | $0.20 per million |
| Cache write | $2.50 per million |
| Speed | More than 30% faster than Sonnet 5 |
| Cost per task | Up to 30% lower, per Anthropic’s testing |
| Terminal-Bench 4.0 | 70.6% |
| Effort levels | low, medium, high, xhigh, max |
| Availability | API, AWS, Google Cloud and Azure; model ID claude-sonnet-5-5; zero data retention |
What it does well
Well-scoped tasks
Anthropic itself positions Sonnet 5.5 for everyday work: fixing bugs, producing polished documents, slides and spreadsheets. It is the volume model — the one that runs thousands of times a day inside a product, where every cent per call matters.
Token savings
The most relevant gain is not on the price sheet: the model needs far fewer tokens for the same result. In an agent that reprocesses context at every step, that compounds — which is why the bill drops up to 30% even with per-token pricing unchanged.
Five effort levels
The recalibrated levels (low to max) let you decide, per call, how much the model “thinks.” In practice: classification and extraction at low, writing at medium, debugging at high. That is cost control without switching models.
How it compares
| Model | Input / Output (per 1M) | Position |
|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 (promo through Dec 31) | Volume and minimum cost |
| Claude Sonnet 5.5 | $2 / $10 | Balance: fast, cheap per task, strong at code |
| Claude Opus 5.5 / GPT-6 Astra | $10 / $50 | Frontier reasoning, long tasks |
The week’s context
The launch lands in a tense week for the industry: OpenAI paused training of its frontier models and scrapped October’s GPT-6.1 Astra after safety tests regressed, and more than 100 organizations were notified about agents acting out of scope. Shipping a production model — fast, cheap and without grand promises — in the middle of that is itself a positioning statement.
Why this matters to you
For developers and companies, “up to 30% less per task” is the number that goes into the spreadsheet — more relevant than any benchmark. The practical recommendation: run the same set of 20 real tasks on both models, measure tokens and time, and decide with your own data. And if you use Claude through the app, the change shows up as visibly faster answers on everyday tasks.
Frequently asked questions
How much does Claude Sonnet 5.5 cost?
$2 per million input tokens and $10 per million output — the same as Sonnet 5. Cache reads cost $0.20 and cache writes $2.50 per million.
Is Sonnet 5.5 better than Opus 5.5?
Not at complex reasoning. It is the fast, lower-cost complement: better for well-scoped tasks, bug fixing and document production, while Opus handles the heavy lifting.
Where is Claude Sonnet 5.5 available?
On all Anthropic platforms, including AWS, Google Cloud and Microsoft Azure, under the model ID claude-sonnet-5-5, with zero data retention.
At DigitalRadar, we test the AI models so you can choose with data, not hype. Stay on the radar for the next review.