Cryptocurrency Prices by Coinlib

Anthropic's Claude Sonnet 5.5 Is Out, Beats Opus 5.5 at Coding for Half the Worth – Decrypt

Briefly
Anthropic launched Claude Sonnet 5.5 on Monday at $2 per million enter tokens and $10 per million output tokens—unchanged from Sonnet 5 and half the worth of Opus 5.5.
Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0 in opposition to 66.4% for Opus 5.5, per Anthropic; unbiased tester Synthetic Evaluation additionally has it forward, 63.6% to 59.6%.
Synthetic Evaluation ranks it second behind Opus 5.5, however says it used extra tokens per activity than any mannequin it examined.
Anthropic launched Claude Sonnet 5.5 on Monday, an improve to Sonnet 5 from June. Anthropic says this middle-tier mannequin runs greater than 30% sooner than its predecessor.“Sonnet 5.5 is strongest at well-scoped on a regular basis duties, fixing bugs, and creating polished paperwork, slides, and spreadsheets. It’s additionally obtained a pointy eye for design,” Anthropic wrote.Myriad: How low will Nvidia go? Click on to make your prediction.The worth stays at $2 per million enter tokens and $10 per million output tokens. Tokens are the chunks of textual content an AI reads and writes, a bit shorter than a phrase, and firms invoice by the million. That's half what Opus 5.5 fees. Nevertheless, it makes use of practically a 3rd much less tokens per activity, which suggests it finally ends up being cheaper to run than Sonnet 5.That mentioned, this mannequin shines in coding. On Terminal-Bench 4.0—a take a look at of whether or not an AI agent can end advanced skilled duties by typing instructions by itself, scored because the share of duties accomplished—Sonnet 5.5 hit 70.6%. Opus 5.5 scored 66.4%, and Sonnet 5 managed 10.3%.In plain phrases, the cheaper mannequin completed extra jobs. Synthetic Evaluation, an unbiased testing agency, ran its personal model and agrees: 63.6% for Sonnet 5.5, 59.6% for Opus 5.5, and 59.1% for OpenAI's GPT-6 Astra.
Claude Sonnet 5.5 (max) makes massive strides on Terminal-Bench, sitting among the many high fashions for each Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 level enhance over Claude Sonnet 5 (max), and barely above 60% for Opus 5.5 and GPT-6… pic.twitter.com/7CPrhgmfxe
— Synthetic Evaluation (@ArtificialAnlys) September 28, 2026Scores additionally depend upon the trouble setting, a dial that makes a mannequin suppose longer for a greater reply and a much bigger invoice. Anthropic says Sonnet 5.5 at Excessive effort matches GPT-6 Sol on FrontierCode for a couple of fifth of the associated fee per activity.On GDPval-AA, which grades real-world skilled work throughout 44 occupations utilizing Elo—the chess-style system that ranks relative ability—Sonnet 5.5 scored 1844 to Opus 5.5's 1846, successfully a tie. GPT-6 Sol scored 1487.Rivals match the worth. OpenAI minimize GPT-6 Sol to $2 and $10 final week, and GPT-5.6 Terra, its mid-tier mannequin, lists at $2 and $12. Anthropic revealed no Terra benchmarks.The catchSonnet 5.5 is a heavy talker. At max effort it wrote about 193,000 tokens per take a look at activity, probably the most Synthetic Evaluation has measured and roughly 60% greater than Opus 5.5. That got here to $7.60 per activity, about 50% above Sonnet 5, which cuts in opposition to Anthropic's declare of as much as 30% financial savings.Anthropic's financial savings come from decrease settings: at Medium effort, the default in its apps, it says Sonnet 5.5 beats Sonnet 5's finest coding rating for lower than a tenth of the associated fee. Synthetic Evaluation says Excessive effort is the most effective worth. For on a regular basis customers, meaning near-flagship coding at a fraction of the worth, so long as the dial stays low.Anthropic's desk is self-reported, and Synthetic Evaluation examined a pre-release construct with a bug that Anthropic expects modified little or barely understated its scores. Anthropic says Opus 5.5 stays clearly stronger at advanced work needing sustained judgment.Claude Haiku 5.5, constructed for high-volume, cost-sensitive purposes, is due within the coming weeks.Each day Debrief NewsletterStart each day with the highest information tales proper now, plus unique options, a podcast, movies and extra.