Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task
Key Points
Anthropic has released Claude Sonnet 5.5, which generates output more than 30 percent faster and cuts per-task costs by up to 30 percent through more efficient token usage.
Sonnet 5.5 shows major gains in coding and knowledge-work benchmarks, nearly matching the top-tier Opus 5.5 in some tests at a fraction of the cost.
The model is available now on AWS, Google Cloud, and Azure. Anthropic has added new safeguards against cybersecurity risks and distillation attacks.
Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on several benchmarks.
Opus 5.5 is built for complex tasks that demand careful judgment. Sonnet 5.5 targets well-defined everyday work like fixing bugs, writing docs, building presentations, and creating spreadsheets. Anthropic also announced Claude Haiku 5.5 for the coming weeks. That model will focus on high-throughput, low-cost use cases.
With Fable, Opus, and Sonnet, Anthropic already has counterparts to OpenAI's GPT-6 Astra, Sol, and Luna , which kicked off the latest pricing battle. Roughly speaking, Opus sits a bit above Sol, Sonnet above Luna, and Fable above Astra, though Anthropic charges more across the board. Performance differences between matched tiers are small enough that cost may end up being the deciding factor. Haiku could help Anthropic close that gap. Ad
Coding performance jumps sharply
The performance gap between Sonnet 5.5 and its predecessor is most significant in coding, according to Anthropic. On Terminal-Bench 4.0, a test for agentic coding, Sonnet 5.5 hits 70.6 percent compared to Sonnet 5's 10.3 percent. On CursorBench 4.0, which recreates real coding sessions from the Cursor editor, Sonnet 5.5 scores 55.5 percent versus 34.1 percent, landing just two points below Opus 5.5 (57.8 percent). Ad
On FrontierCode 1.1 at the "High" setting, Sonnet 5.5 scores ten points above Sonnet 5 at roughly one-fifteenth the cost per task, Anthropic says. Early testers praised how quickly the model grasps a codebase. Sonnet 5.5 also batches tool calls more often than its predecessor, which cuts the number of steps needed.
Reasoning intensity has one odd wrinkle. At the highest effort level, "Max," Sonnet 5.5 actually scores worse on FrontierCode than at "Xhigh." Anthropic says that at maximum effort, the model more frequently triggers a code-review function that splits work across multiple sub-agents. In some cases, that led to timeouts or changes outside the task scope, both of which FrontierCode penalizes. Ad
Knowledge work nearly matches Opus 5.5
On GDPval-AA, an OpenAI-developed knowledge-work benchmark covering tasks from 44 professions and nine industries, Sonnet 5.5 scores 1,844 points. That nearly matches Opus 5.5 (1,846) and sits about 400 points above Sonnet 5 (1,449). OpenAI's GPT-6 Sol lands at 1,487 by Anthropic's numbers. On Chartography, a visual chart recognition test, Sonnet 5.5 jumps from 15.6 to 61.6 percent.
¹ For Opus 5.5, Terminal-Bench 4.0 shows the highest score at the "Xhigh" setting. ² Sonnet 5.5 scores lower on FrontierCode at "Max" than at "Xhigh." ³ Sonnet 5.5 values come from a pre-release version where a since-fixed bug may have affected structured outputs. Ad
Early testers described Sonnet 5.5 as a more natural conversational partner with a feel for design. The model can rework user interfaces and execute slide templates so well that results barely need touch-ups, Anthropic claims. Ad
Anthropic also says Sonnet 5.5 is the first Sonnet model that can play through Pokémon Red using only screenshots. OpenAI's Astra recently showed similar progress on gaming benchmarks . Tasks like these were considered hard just a few years ago and often required custom algorithms. Now even mid-tier models within a product family can handle them.
Same token price, fewer tokens per task
Sonnet 5.5 costs the same per million tokens as Sonnet 5: $2 for input tokens, $10 for output tokens, and $0.20 for cache reads. Because the model uses fewer tokens per task, effective costs drop by up to 30 percent, Anthropic says. Output generation is also more than 30 percent faster. Independent testing still needs to confirm these claims.
Like Anthropic's other models and OpenAI's counterparts, Sonnet 5.5 offers an adjustable effort setting that lets users trade cost and speed against output quality. At low or medium settings, Sonnet 5.5 already beats Sonnet 5's best scores on several benchmarks at about one-tenth the per-task cost, Anthropic says. Finding the right effort level for each task, though, is still more art than science.
New cybersecurity safeguards for a more capable model
Claude Sonnet 5.5 is available now across all major cloud platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Like Opus 5.5 and Sonnet 5, Anthropic offers the model with zero data retention. Developers can access it through the Claude Platform using the model ID "claude-sonnet-5-5."
Because Sonnet 5.5 is far more capable in cybersecurity than its predecessor, Anthropic is adding safeguards to a Sonnet model for the first time. Requests involving high-risk cybersecurity tasks get visibly rerouted to Sonnet 5. Through an expanded Cyber Verification Program , qualified professionals can apply for tiered access.
Anthropic has also added safety classifiers against distillation attacks, the same kind used in its most powerful models. Whether these measures actually work should become clear over the coming months. If Chinese labs have been benefiting from such attacks , closing that vector could widen the gap with Western labs again.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.