All insights
AI in Marketing
8 min read28 September 2026Nathan Mzumara

Claude Sonnet 5.5 Holds Its Price, Cuts Cost Per Task 30%

Claude Sonnet 5.5 Holds Its Price, Cuts Cost Per Task 30%

Anthropic released Claude Sonnet 5.5 on 28 September 2026, the second model in the Claude 5.5 family. The price per token is unchanged from Claude Sonnet 5, yet Anthropic says the model runs 30%+ faster and costs up to 30% less for most work, because it typically needs far fewer tokens to finish the same job.

What Anthropic shipped on 28 September 2026

Anthropic announced Claude Sonnet 5.5 on 28 September 2026 as the second model in the Claude 5.5 family. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the family in the coming weeks.

Pricing is $2 per million input tokens, $10 per million output tokens and $0.20 per million tokens for cache reads. Those are the same three numbers Sonnet 5 carried. Medium effort is the default in the Claude apps. High effort is the default on the Claude Platform.

Anthropic positions the new model as a faster, lower-cost complement to Claude Opus 5.5. Opus 5.5 is built for complex work requiring careful judgment. Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets, and its speed suits fast iteration on less complex tasks. Full details sit in Anthropic's Claude Sonnet 5.5 announcement.

Why does the cost fall if the price per token has not changed?

Claude Sonnet 5.5 is a mid-tier large language model from Anthropic that completes the same work as Claude Sonnet 5 using far fewer tokens, so the cost of a finished task falls while the list price stays flat. That is the whole story in one sentence.

Anthropic states the model typically needs far fewer tokens to do the same work, and that in its own testing it costs up to 30% less per task than its predecessor. No line on the rate card moved.

Anthropic also reports that early testers were struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs. Fewer steps is the mechanism. Lower spend is the result.

Effort level is now part of the price

Anthropic reports that on several evaluations, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. It says the model complements Opus 5.5 best at lower effort settings, where it costs less per task, and that at higher settings it can perform comparably at a similar cost.

For anyone routing work between models, that reframes the escalation decision. Turning effort up on the cheaper model is now a live alternative to jumping to the flagship.

How does the new model score against Sonnet 5 and Opus 5.5?

The headline result is agentic coding. Anthropic reports 70.6% on Terminal-Bench 4.0 for Sonnet 5.5, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5.

EvaluationSonnet 5.5Sonnet 5Opus 5.5
Terminal-Bench 4.070.6%10.3%66.4%
FrontierCode 1.1 (Main)46.2% Max, 52.1% Xhigh42.4%54.4%
CursorBench 4.055.5%34.1%57.8%
GDPval-AA v2.1184414491846
AA-Briefcase v1.1181113591822
Humanity's Last Exam (with tools)64.5%54.9%67.7%
OSWorld 2.1 (partial)80.1%57.0%81.8%
Chartography (no tools)61.6%15.6%64.4%

All figures are Anthropic's, published 28 September 2026. On GDPval-AA v2.1, a test of real-world work across a variety of occupations, Sonnet 5.5 lands two points below Opus 5.5 and roughly 400 points above Sonnet 5. On AA-Briefcase v1.1 the gap to Opus 5.5 is eleven points.

Anthropic also reports that Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots, which it cites as evidence of long-horizon work and image understanding. Method notes are in the Sonnet 5.5 system card.

Anthropic is explicit about the limit of all this. It states that benchmark scores capture only one facet of a model's capabilities, and that in its own testing and that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

Infographic on Anthropic's Claude Sonnet 5.5, released 28 September 2026: up to 30% lower cost per task at unchanged prices, with benchmark and pricing comparisons.
Infographic summary of this articleDownload infographic

What this changes for growth, search and content teams

If you run content pipelines, entity extraction, SERP parsing, AI-visibility monitoring or agentic reporting at scale, this release changes your planning unit rather than your rate card.

  • Cost per million tokens stops being a budget input: the price is identical, so any saving or overspend now comes from how many tokens a model burns to finish a job. Two models at the same rate can produce very different invoices.
  • Cost per completed task becomes the metric: price a crawl parsed, a brief produced, a citation audit run. Anthropic's own charts plot score against cost per task at every effort level, which tells you where it expects the comparison to happen.
  • Effort level belongs in your cost model: Medium is the default in the Claude apps and High is the default on the Claude Platform. Teams running the same prompts in both places are already paying two different amounts per task without changing a setting.
  • Speed buys you volume, not just patience: 30%+ faster generation, which Anthropic calls its fastest Sonnet model to date, shortens iteration loops on daily reporting and monitoring jobs.
  • Cache reads are the cheapest lever you already own: at $0.20 per million tokens, prompt caching still does more for a repetitive monitoring workload than any model swap.

My observation from running these pipelines: most teams still forecast AI spend from a rate card in a spreadsheet, and almost none instrument token consumption per completed job. That gap is now the difference between a 30% saving and a surprise.

The precedent: reasoning models moved the other way

This is the inverse of what happened when reasoning models arrived. Headline per-token prices came down, invoices went up, because thinking tokens were invisible in the pricing table and nobody had modelled them.

Claude Sonnet 5.5 flips the same mechanism in the buyer's favour. Flat sticker price, lower real spend, and no way to verify either from the rate card alone. The lesson is identical in both directions: the pricing table is not the price. I made a related argument about why the spread in LLM API pricing now matters more than the average, and this release is that argument in practice.

It also echoes a pattern seen elsewhere this year, where a new version lands at the old version's prices with higher scores, as happened when Grok 4.7 arrived at Grok 4.6 prices.

Safeguards: the first Sonnet to ship with frontier-level cyber protections

Anthropic reports that because the cybersecurity capabilities of Sonnet 5.5 are comparable to Opus 5's, it is the first Sonnet model to launch with cyber safeguards and fallbacks like those developed for Anthropic's most capable models. Its biology safeguards are the same as Sonnet 5's.

Anthropic states both safeguards target a narrow set of high-risk requests, and that routine software development and most life sciences work are unaffected. For marketing workloads the practical impact is close to zero. The signal matters more than the restriction: mid-tier models now carry frontier-level risk profiles, which will eventually show up in procurement and security reviews.

What could go wrong

"Up to 30% less" is a ceiling measured in Anthropic's testing, not a floor guaranteed on your workloads. A saving on agentic coding tasks does not automatically transfer to a long-context SERP parsing job with heavy input and short output.

There is also a quiet trap in the defaults. High effort on the Claude Platform means longer runs and higher cost per task than the Medium default in the Claude apps. A team that migrates from app testing to API production could watch spend rise while believing it just switched to a cheaper model.

The counter-view

The case against moving quickly is straightforward. Anthropic itself says Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment, and the FrontierCode 1.1 and AA-Briefcase v1.1 gaps back that up. Strategy work, positioning and anything that needs a defensible argument should stay where judgment is best.

Claude Haiku 5.5 is also due in the coming weeks, aimed at high-volume and cost-sensitive applications. If your pipeline is mostly high-volume classification and extraction, a migration decided this month may need revisiting next month.

What to do with Claude Sonnet 5.5 this week

  1. Pick your five highest-volume AI jobs and log tokens in, tokens out, cache reads and wall-clock time per completed task. Without that baseline you cannot prove a 30% saving or spot a 30% rise.
  2. Re-run those five jobs on Sonnet 5.5, Sonnet 5 and Opus 5.5 at matched effort levels, then compare cost per finished output rather than cost per million tokens.
  3. Test raising effort on the cheaper model before escalating to Opus 5.5. Anthropic reports comparable performance at higher settings for similar cost on several evaluations.
  4. Audit your effort defaults across apps and API so production is not silently running at High when Medium would do.
  5. Rewrite your AI budget line from cost per million tokens to cost per completed task, and hold your next quarterly forecast to that unit.

Claude Sonnet 5.5 has not made AI cheaper on paper. It has made the paper unreliable. Teams that instrument task-level telemetry this quarter will know exactly what they are spending. Everyone else will be guessing from a rate card that has not moved since Sonnet 5.

This Claude Sonnet 5.5 analysis was published 28 September 2026. Last updated 28 September 2026.

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.