Marktkapitalisierung —24h Volumen —BTC —Angstindex —
MasterInvestierenInvestitionen & Einnahmen
Anzeige
Anzeige

Anthropic's New Sonnet Model Nears Opus Performance but Uses More Tokens

lesen190AnteilDruckversion
Anthropic's New Sonnet Model Nears Opus Performance but Uses More Tokens

Anthropic released Claude Sonnet 5.5 on September 28 at the same list price as Sonnet 5. The company says the model runs over 30% faster and costs up to 30% less per task.

An independent benchmarking firm reached a different conclusion on cost once it pushed the model to its limit.

Follow us on X to get the latest news as it happens

Anthropic’s Second 5.5 Model Lands Days After Opus

Sonnet 5.5 is the second model in the Claude 5.5 family. Anthropic launched Opus 5.5 on September 22. On the same day, OpenAI also released GPT-6 Sol and Luna.

The new model keeps Sonnet 5’s prices of $2 per million input tokens and $10 per million output tokens. Anthropic reported a 70.6% score on Terminal-Bench 4.0, an agentic coding test. Sonnet 5 managed just 10.3%, while Opus 5.5 reached 66.4%.

“Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets,” the team said.

Anthropic said Sonnet 5.5 does not extend its capability frontier. Its alignment review therefore focused on risks such as misleading users and aiding high-stakes misuse.

Sonnet 5.5 also ships with cyber safeguards and anti-distillation classifiers. These block attempts to extract the model’s reasoning for training rival systems. Claude Haiku 5.5 follows in the coming weeks.

Meanwhile, OpenAI shelved an upcoming model, GPT-6.1 Astra, on safety grounds. CNBC confirmed on Monday that the model fell short of the company’s standards.

Artificial Analysis Finds a Heavier Token Bill at Max Effort

Artificial Analysis, the benchmarking firm that ranked Grok 4.7 fourth, gave Sonnet 5.5 a 56 on its Intelligence Index. That puts it second overall, 2 points behind Opus 5.5 at maximum effort.

The firm recorded 64% for Sonnet 5.5 on Terminal-Bench 4.0, against 60% for Opus 5.5 and GPT-6 Astra. On knowledge-work tests such as GDPval-AA, it found Sonnet 5.5 roughly level with Opus 5.5.

The firm said Sonnet 5.5 used significantly more tokens to reach those results.

“At max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task,”

That volume sits about 60% above Opus 5.5 and Sonnet 5 at max effort. It is also roughly 7 times GPT-6 Astra’s max-effort count. 

Early-access customers quoted by Anthropic reported the opposite trend against Sonnet 5, though on different workloads. Balyasny Asset Management saw far lower token use on its finance tasks.

“On our private suite of 2,441 finance tasks covering Q&A, extraction, analysis, and forecasting, Claude Sonnet 5.5 scored ahead of Sonnet 5 and used about 121k tokens per answer, where Sonnet 5 used 497k,”  Joe Poirier, Senior AI Engineer at Balyasny Asset Management, stated.

Artificial Analysis tested a pre-release build carrying a structured-output bug that Anthropic has since fixed. The firm plans to rerun the relevant evaluations soon.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Source: BeInCrypto

Mehr zum Thema «Cryptocurrency News»

Alle Beiträge
Ein zufälliges Zitat über Geld
Люди со средствами думают, что главное в жизни - любовь; бедняки знают точно, что главное - деньги.
— Джералд Бренан

Interessant in anderen Abschnitten

Ganzen Blog

Kommentare 0

Noch keine Kommentare

Seien Sie der Erste, der Ihre Meinung oder Erfahrung zu diesem Thema teilt.

Anzeige