Claude Sonnet 5.5 outscores Opus 5.5 on Terminal-Bench, and gets worse at max effort
Anthropic released Claude Sonnet 5.5 today, Monday, Sep 28, six days after Opus 5.5 was released. In the Antropic announcement, it says it improves clearly on Sonnet 5, generates output more than 30% faster, and costs up to 30% less per task.
Anthropic describes Sonnet 5.5 as a model that comes close to Opus. Its own benchmark table shows one case where it beats Opus.
On Terminal-Bench 4.0, an agentic coding test run in a command line, Sonnet 5.5 scored 70.6%. Opus 5.5 scored 66.4%, and a footnote says that figure is Opus at Xhigh effort, its best result. The cheaper model finished ahead of the flagship on the flagship’s best run. Anthropic still says Opus 5.5 is clearly stronger at complex, open-ended work that requires sustained judgment, and that one benchmark doesn’t overturn that.
Developers who run terminal agents will notice the result.
Editor's Picks
More effort, worse score
On FrontierCode, which checks whether an agent’s code changes would be merged as-is, Sonnet 5.5 scored 52.1% at Xhigh effort and 46.2% at Max. Adding effort made it worse.
Anthropic explains why in a footnote. At Max, Sonnet 5.5 more often ran Claude Code’s code-review skill, which splits a review across many subagents. In two cases examined by Cognition, that led to a timeout or to edits outside the task’s scope, and FrontierCode penalizes out-of-scope changes even when they are good.
If you set the model to maximum effort on the assumption that more thinking always helps, you may end up with changes you didn’t ask for.
Half price, except for cache reads
Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5 and exactly half of Opus 5.5’s $4 and $20. Cache writes are also half, at $2.50 versus $5. Cache reads are $0.20 per million for both models.
Agent workloads reuse large contexts over and over, so cache reads are a big part of their bills. For those workloads, Sonnet saves less than “half the price of Opus” implies, because the line item that dominates is priced the same.
The token savings are larger than the headline
Anthropic’s “up to 30% less per task” is modest compared with some customer results. Balyasny Asset Management ran 2,441 finance tasks and reported about 121,000 tokens per answer from Sonnet 5.5, versus 497,000 from Sonnet 5.
That is roughly a quarter of the tokens.
Box reported 12% fewer total tokens, and Slack about 14% fewer output tokens. Savings depend heavily on the workload, so the only reliable number for your company is the one you measure on your own tasks.
Check this before you migrate
Anyone who runs Sonnet with thinking turned off has to switch to a new between tools setting before moving to Sonnet 5.5, according to Anthropic’s migration guide.
As protection against distillation, Sonnet 5.5 expands preserved thinking so that Claude’s reasoning stays tied to the account that generated it. If you switch accounts in the middle of a Claude Code session, this will definitely affect you.
Security work is also restricted.
Anthropic rates Sonnet 5.5’s cyber capabilities as a large step up from Sonnet 5, so higher-risk cybersecurity requests will visibly fall back to Sonnet 5. Routine bug fixing is not affected.
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work. pic.twitter.com/UvXD8mDTF1
— Claude (@claudeai) September 28, 2026
The watermark deadline is Wednesday
Every Claude model released since early September carries Anthropic’s invisible text watermark from day one, and users cannot turn it off. Anthropic’s version is based on Google DeepMind’s SynthID-Text. Instead of adding hidden characters, it slightly shifts which of several plausible next words the model picks, creating a statistical pattern that a key can detect across longer passages.
According to a September 24 customer email reported by The Decoder, Sonnet 5, Opus 4.8 and Fable 5 start producing watermarked output on September 30. That means Sonnet 5, the model Sonnet 5.5 falls back to for sensitive cyber requests, becomes watermarked two days after this launch.
Anthropic says Sonnet 5.5 writes more clearly and in a more human way. Its text is also designed to be detectable as AI-generated. Anthropic itself notes that a positive detection only indicates Claude was likely involved; it could have translated, summarized or edited text someone else wrote.
What stays unresolved
Some of the benchmarks are under revision. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a bug affecting structured outputs. Anthropic says the effect was probably small and has been fixed. On GDPval-AA, Sonnet 5.5 scored 1844 to Opus 5.5’s 1846, a gap small enough that the bug could plausibly account for it in either direction.
Terminal-Bench and CursorBench haven’t published scores for GPT-6 Sol, so those charts compare against GPT-5.6 Sol.


Join the discussion
Be the first to join this discussion
Be respectful toward authors and fellow readers.
Report comment
Tell the moderation team why this comment may break the community guidelines.