top of page

Claude Sonnet 5: Price-Performance, Office Work, Coding, and Everyday Business Use

1 day ago
6 min read

Claude Sonnet 5 is Anthropic's mid-tier model, launched June 30, 2026, built to close most of the gap to the flagship Opus tier while staying at Sonnet-level pricing.

........

  • Sonnet 5 sits between Haiku 4.5 and Opus 5 in Anthropic's current lineup, and replaced Sonnet 4.6 as the default model for Free and Pro plans on launch day.

  • It runs $2 per million input tokens and $10 per million output tokens — a rate Anthropic made permanent on August 10, 2026, cancelling a planned increase to $3/$15 that had been scheduled for September 1.

  • Sonnet 5 uses a new tokenizer, the same one introduced with Opus 4.7, which turns the same text into roughly 1.0 to 1.35 times more tokens than Sonnet 4.6 produced.

  • On SWE-bench Verified it scores 72.7%, up from Sonnet 4.6's 62.3%, and closing in on Opus 4.8's 79.4%.

  • It supports a 1 million token context window, up to 128,000 tokens of output, and five selectable effort levels for reasoning depth.

··········

WHERE SONNET 5 SITS IN THE LINEUP.

Sonnet 5 is positioned as the mid-tier option between Haiku, built for speed and low cost, and Opus, built for maximum capability.

At launch, Anthropic framed Sonnet 5 as a model that closes most of the performance gap to Opus-class models while keeping Sonnet-level pricing.

The comparison at launch was against Opus 4.8, which was Anthropic's flagship at the time.

Opus 5 has since shipped, becoming the new flagship, but Sonnet 5's position in the lineup — the mid-tier, cost-efficient option — hasn't changed.

Starting the Tuesday of its launch, Sonnet 5 became the default model for both Free and Pro plans on Claude.ai, and it's also selectable for Max, Team, and Enterprise users, inside Claude Code, and on the Claude Platform.

··········

BENCHMARK RESULTS.

Sonnet 5 beats its predecessor, Sonnet 4.6, on every benchmark Anthropic published at launch, and narrows the distance to Opus 4.8 on most of them.

On SWE-bench Verified, a coding benchmark, Sonnet 5 scores 72.7%, compared with 62.3% for Sonnet 4.6 and 79.4% for Opus 4.8.

On Terminal-bench, which measures agentic command-line task completion, Sonnet 5 jumps to 76.1%, up from Sonnet 4.6's 55.4% — a 20.7 point gain, the single largest improvement Anthropic reported at launch.

On SWE-bench Pro, a harder agentic coding benchmark, Sonnet 5 scores 63.2%, against Sonnet 4.6's 58.1% and Opus 4.8's 69.2%.

On OSWorld-Verified, a computer-use benchmark, Sonnet 5 reaches 81.2%.

On Humanity's Last Exam, a graduate-level multidisciplinary reasoning benchmark, Sonnet 5 scores 57.4%, which several independent write-ups noted actually edges out Opus 4.8 on knowledge-work-style questions specifically, even though Opus remains stronger on the hardest reasoning problems overall.

··········

Sonnet 5 vs. Sonnet 4.6 vs. Opus 4.8: key benchmarks

Benchmark

Sonnet 4.6

Sonnet 5

Opus 4.8

SWE-bench Verified

62.3%

72.7%

79.4%

Terminal-bench

55.4%

76.1%

SWE-bench Pro

58.1%

63.2%

69.2%

OSWorld-Verified

81.2%

Humanity's Last Exam

57.4%

··········

PRICING, AND THE CANCELLED SEPTEMBER INCREASE.

Sonnet 5 launched at $2 per million input tokens and $10 per million output tokens, explicitly framed as introductory pricing through August 31, 2026.

The plan at launch was for pricing to rise to $3 per million input tokens and $15 per million output tokens starting September 1 — a 50% increase across the board.

On August 10, 2026, Anthropic reversed that plan. The company's own announcement stated it plainly: the $2/$10 pricing is now the standard price, and the scheduled September 1 increase will not occur.

That makes Sonnet 5 unusually positioned: a newer-generation model priced lower than its own predecessor, since Sonnet 4.6 sits at $3 per million input tokens and $15 per million output tokens.

It also undercuts Opus 4.8, which runs $5 per million input tokens and $25 per million output tokens, and compares favorably with GPT-5.6 Sol and Gemini 3.1 Pro on a pure per-token basis, according to reporting at launch.

··········

Sonnet 5 pricing: what was planned vs. what happened


Planned (at launch)

Actual (from August 10, 2026)

Input, through Aug 31

$2 / million tokens

$2 / million tokens

Output, through Aug 31

$10 / million tokens

$10 / million tokens

Input, from Sept 1

$3 / million tokens

$2 / million tokens (unchanged)

Output, from Sept 1

$15 / million tokens

$10 / million tokens (unchanged)

··········

THE TOKENIZER CATCH.

The permanent $2/$10 rate is not the whole cost picture, because Sonnet 5 counts tokens differently than Sonnet 4.6 did.

Sonnet 5 uses the tokenizer Anthropic first introduced with Opus 4.7, rather than the tokenizer Sonnet 4.6 and earlier models used.

For the same piece of text, that newer tokenizer produces roughly 1.0 to 1.35 times more tokens, depending on the content and workload shape.

In practice, that means a straight per-token price comparison between Sonnet 5 and Sonnet 4.6 can overstate the savings: the rate per token dropped, but the number of tokens billed for the same job may have gone up by as much as a third.

Anthropic has said it raised rate limits to accommodate the higher token counts that come with running Sonnet 5 at higher effort levels.

··········

REASONING EFFORT AND SPECS.

Sonnet 5 supports adaptive thinking, with a selectable effort parameter that trades cost for reasoning depth.

Effort levels run low, medium, high, max, and x-high, letting developers choose how much the model reasons before answering, on a per-request basis.

The context window is 1 million tokens, with a maximum output of 128,000 tokens per response.

Sonnet 5 accepts text, image, and file inputs, and includes real-time cyber safeguards that block certain high-risk dual-use activity.

One tradeoff worth noting: at the highest effort setting, x-high, Sonnet 5 can end up costing more per task than Opus 4.8 for comparable quality, since heavier reasoning burns through tokens fast even at the lower per-token rate. Sonnet 5's value case is strongest at low and medium effort, where most everyday and business tasks sit.

··········

OFFICE WORK AND EVERYDAY BUSINESS USE.

Beyond coding benchmarks, Sonnet 5 is positioned by Anthropic and by third-party API resellers as a general business workhorse.

Typical use cases cited include document drafting, spreadsheet analysis, presentation building, and other structured office tasks that don't require Opus-level reasoning but benefit from more reliability than a cheaper Haiku-class model offers.

Production agent deployments are the other big use case: high-volume pipelines where a task gets repeated thousands of times, and small per-token savings compound into a meaningful budget difference.

Anthropic's own framing at launch emphasized agentic reliability over any single headline benchmark — longer task chains without losing context, better recovery when a tool call fails, and steadier behavior across long sessions in Claude Code or Cowork.

··········

WHAT EARLY USERS REPORTED.

A handful of Anthropic's launch partners described specific, concrete changes in how their workflows ran on Sonnet 5.

Cursor reported agents that stayed on plan and shipped clean multi-step code changes at an efficient cost, rather than drifting off-task partway through longer jobs.

Lovable highlighted cleaner refusals of unsafe requests compared with the previous generation.

ClickHouse pointed to tighter reasoning steps and faster time to insight on its internal workloads.

One tester's description that circulated after launch: Sonnet 5 writing a reproducing test for a bug, fixing the bug, then re-running the original test to confirm the fix actually held — all in a single pass, without being prompted step by step.

··········

HOW IT COMPARES TO GPT-5.6 TERRA.

On raw coding benchmark scores, Sonnet 5 and OpenAI's GPT-5.6 Terra land close together.

Sonnet 5 scores 63.2% on SWE-bench Pro, against GPT-5.6 Terra's reported 63.4% on the same benchmark — a gap small enough to fall inside normal run-to-run variance.

Both figures are vendor-reported, and OpenAI itself published an audit in July 2026 estimating that roughly 30% of SWE-bench Pro tasks contain flawed test criteria or misleading task descriptions, which is a reason to treat any single benchmark score on that suite with some caution.

For most teams, cost per completed task on their own actual workload is a more reliable comparison than a sticker price or a single leaderboard number.

··········

AVAILABILITY AND ACCESS.

Claude Sonnet 5 is available today across Claude.ai, the Claude API, Claude Code, and the Claude Platform.

It's the default model for Free and Pro plans on Claude.ai, and it's selectable for Max, Team, and Enterprise users.

Developers can call it through the Claude API using the model string claude-sonnet-5, and it's also accessible through third-party routers and OpenAI-compatible API proxies that host Anthropic models alongside others.

··········

·····

FOLLOW US FOR MORE.

·····

·····

DATA STUDIOS

·····

Recent Posts

See All
bottom of page