top of page

Claude Sonnet 5 vs GPT-5.6 Sol vs Gemini 3.8 Flash: Complete Comparison and Report on Pricing, Benchmarks, Context Window, and Value-Tier Positioning

8 minutes ago
7 min read

Three models currently anchor the value tier at the three largest labs: Claude Sonnet 5, launched by Anthropic on June 30, 2026; GPT-5.6 Sol, launched by OpenAI on July 9, 2026; and Gemini 3.8 Flash, launched by Google on September 2, 2026. Unlike a flagship-versus-flagship comparison, these three were built to compete directly on cost per completed task, and each vendor explicitly frames its release around getting near-frontier results at a fraction of frontier pricing.


This report also has to flag something upfront: pricing on two of these three models has moved since launch, in opposite directions, and getting the current rate wrong changes the entire comparison.


··········


RELEASE TIMELINE AND MODEL IDENTITY.

Model identifiers, lineage, and how each vendor positions its value tier.


........


Attribute

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

Vendor

Anthropic

OpenAI

Google DeepMind

Release date

June 30, 2026

July 9, 2026 (limited preview June 26)

September 2, 2026

Model ID

claude-sonnet-5

gpt-5.6-sol

gemini-3.8-flash

Predecessor

Claude Sonnet 4.6

GPT-5.5

Gemini 3.7 Flash

Position in lineup

Mid-tier, between Haiku and Opus

Flagship of the GPT-5.6 family (Sol/Terra/Luna)

Mid-tier, below the Pro line

Base model status

New generation

New generation, distilled with Terra and Luna from one training run

Built on the Gemini 3.7 Flash architecture


........


Only Sol carries flagship status inside its own family; Sonnet 5 and Gemini 3.8 Flash are both deliberately positioned as the economical option below a more expensive sibling tier at their respective labs. That distinction matters for reading the benchmark tables later in this report: a value-tier model losing to a frontier flagship in another vendor's lineup is not surprising, but a value-tier model beating a same-week flagship on a shared benchmark, which happens more than once below, is worth noting explicitly.


··········


PRICING AND THE TWO RATE CHANGES SINCE LAUNCH.

Current confirmed rates, and the corrections needed on top of each launch announcement.


........


Rate per 1M tokens

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

Standard input (current)

$2

$5

$0.75 (through Dec 31, 2026)

Standard output (current)

$10

$30

$3.75 (through Dec 31, 2026)

Cached input

$0.20

$0.50

Not published in this table

Promotional adjustment

Planned increase to $3/$15 on September 1, 2026, since cancelled

Cut by more than 20% from August 21, 2026, through at least November 21, 2026

Doubles to $1.50/$7.50 on January 1, 2027


........


Sonnet 5 launched with a stated introductory period ending August 31, 2026, after which the rate was due to rise to $3 input and $15 output. Anthropic's own pricing page, rechecked independently on August 20 and again on August 26, still showed $2/$10 with no introductory label and no announced increase, and the step-up has since been described as cancelled. Anyone quoting Sonnet 5 at $3/$15 today is citing a change that did not happen.


Sol moved in the opposite direction: OpenAI cut its listed $5/$30 rate by more than 20% starting August 21, 2026, confirmed available through at least November 21. The exact discounted figure varies slightly across sources reporting the cut, but the direction and the 20%-plus magnitude are consistent across OpenAI's own pricing update.


Gemini 3.8 Flash is the only one of the three with a scheduled future increase rather than a past correction: its entire rate card doubles on January 1, 2027, a date that applies across Google's whole Flash 3.x family rather than to this release specifically.


At current confirmed rates, Gemini 3.8 Flash costs roughly a third of Sonnet 5's rate and about a fifteenth of Sol's discounted rate on input, with a comparable multiple on output.


··········


CONTEXT WINDOW, OUTPUT LIMITS, AND TOKENIZER EFFECTS.

Window sizes and a cost variable that does not show up in the headline price.


........


Limit

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

Context window

1,000,000 tokens

1,050,000 tokens

1,048,576 tokens

Maximum output

128,000 tokens

128,000 tokens

65,536 tokens

Knowledge cutoff

Not separately published for Sonnet 5 in this comparison

February 16, 2026

March 2026 for most domains, January 2025 for some

Long-context billing threshold

None

Above 272,000 input tokens: 2x input, 1.5x output, 1.25x cache write

None


........


Gemini 3.8 Flash's output ceiling is exactly half of the other two, a constraint documented in Google's own model card for agent work that reasons extensively before producing final output. Sol carries the same 272,000-token long-context penalty structure this outlet has documented on its flagship sibling Astra, applied here to the value tier as well.


Sonnet 5 introduced an updated tokenizer at launch, the same generation Anthropic later used for Opus 4.7 and Fable 5.1, which maps the same text to roughly 1.0 to 1.35 times more tokens depending on content type than Sonnet 4.6's tokenizer produced. Anthropic states it set the introductory price specifically to make the move from Sonnet 4.6 roughly cost-neutral despite the token-count increase, which means a raw per-token price comparison against Sonnet 4.6, or against any model using an older tokenizer, understates Sonnet 5's real cost per unit of text.


··········


BENCHMARK METHODOLOGY AND SOURCE RELIABILITY.

Why several of the most-cited figures for these three models need a version or date check before use.


None of the three vendors tested its value-tier model against both of the other two under matching conditions at launch. Sonnet 5's system card compares it against GPT-5.5 and Gemini 3.5 Flash, since GPT-5.6 Sol had not been released as of June 30, 2026; any Sonnet-5-versus-Sol comparison circulating today is a later, unofficial pairing rather than a vendor-published one.


Sol's own benchmark reporting is not fully consistent across sources checked for this report. One tracker lists Sol's Terminal-Bench 2.1 score at 88.80%, ranked first among tracked models on that specific benchmark, while a separate outlet's chart compares Sol against unnamed competitors without a table, and a third source's launch-day coverage frames a different model's Terminal-Bench 2.1 figure (83.4% for GPT-5.5) in a way that could be mistaken for Sol's own score if read quickly. The 88.80% figure is treated as Sol's in the table below because it is sourced to a benchmark-specific tracker rather than a comparative launch chart.


Gemini 3.8 Flash's DeepSWE score carries the same version ambiguity documented in this outlet's earlier Astra-versus-Gemini report: the model card cites version 1 of the benchmark, while Google's launch post cites version 1.1, treating the two as equivalent.


··········


DIRECTLY COMPARABLE SCORES.

The one benchmark where all three models have a same-version, independently trackable figure.


........


Benchmark

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

Terminal-Bench 2.1

80.4% (beats Opus 4.8's 74.6% on this specific benchmark)

88.80%

90.8% (up from 81.6% for Gemini 3.7 Flash)

SWE-bench Pro

63.2%

Not published in this comparison

Not published in this comparison

OSWorld-Verified / OSWorld 2.0

81.2%

Not published in this comparison

Not published in this comparison

BrowseComp (agentic search)

84.7%

Not published in this comparison

Not published in this comparison

GDPval-AA v2 (Elo)

1,618

Not published in this comparison

Not published in this comparison

FrontierMath v2

Not published in this comparison

89.12%

Not published in this comparison

Context Arena

Not published in this comparison

97.63%

Not published in this comparison


........


Terminal-Bench 2.1 is the only row with a plausible same-version reading across all three vendors, and on it, the two cheapest models in this comparison outscore the most expensive one: Gemini 3.8 Flash at 90.8% and Sol at 88.80% both beat Sonnet 5's 80.4%, despite Sonnet 5 costing more per token than Gemini and, until August, roughly the same as Sol's original list price. Every other row in this table has data from only one vendor, which means no ranking beyond this single benchmark can currently be supported across all three models.


··········


LATENCY AND THROUGHPUT.

Independently measured speed where available, against undisclosed figures where it is not.


........


Measure

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

p95 time to first token

Not published in this comparison

8.37 seconds

13.30 seconds

Output throughput rank

Not published in this comparison

Not published in this comparison

3rd of 196 models evaluated

Median TTFT across evaluated models

Not published in this comparison

Not published in this comparison

2.99 seconds


........


Sol's 8.37-second time to first token sits closer to the cross-model median than Gemini 3.8 Flash's 13.30 seconds, though the two figures come from different measurement sources and are not guaranteed to use identical methodology. No public latency figure was located for Sonnet 5 in this research pass, which is itself worth noting given how central latency is to a model marketed for high-volume, everyday use.


··········


AVAILABILITY ACROSS PLATFORMS.

Where each model actually runs today.


........


Channel

Claude Sonnet 5

GPT-5.6 Sol

Gemini 3.8 Flash

Default consumer tier

Free and Pro plans on claude.ai

Plus, Pro, Business, Enterprise in ChatGPT

Not specified as a ChatGPT/Claude-style consumer default

Coding agent

Claude Code

Codex

Not listed

Third-party IDE integration

Cursor, VS Code, GitHub Copilot

Not specified

Not specified

Cloud platforms

Claude Platform, Amazon Bedrock, Google Vertex AI, Microsoft Foundry

OpenAI API, Amazon Bedrock, Azure

Google Cloud

Enterprise tiers

Max, Team, Enterprise

Business, Enterprise

Not specified in this comparison


........


Sonnet 5 has the broadest documented third-party IDE presence of the three, appearing natively in Cursor, VS Code, and GitHub Copilot at launch. Sol and Gemini 3.8 Flash are each available on their vendor's own cloud plus at least one additional major cloud provider, while Sonnet 5 is the only one of the three confirmed on all three major clouds.


··········


WORKLOAD ALLOCATION CRITERIA.

What the corrected pricing and the single comparable benchmark actually support as a decision.


The pricing corrections in this report matter more than any single benchmark gap: at current rates, Gemini 3.8 Flash costs roughly a third of Sonnet 5 and a fifteenth of discounted Sol on input tokens, and on the one benchmark where all three can be read against the same version, it does not trail either competitor. That combination makes it the reasonable default to test first for high-volume, cost-sensitive agent work, provided a team accounts for its 65,536-token output ceiling and confirms the January 2027 price doubling does not break the economics of a long-term commitment made today.


Sol's case rests on the benchmarks it publishes and the other two do not: FrontierMath v2 and Context Arena figures with no equivalent Sonnet 5 or Gemini 3.8 Flash score to compare against, plus its 272,000-token pricing threshold, which needs checking against actual prompt-length distribution before it changes the cost picture. Its post-discount price is the one to budget against, not its original $5/$30 list rate, and that discount carries a stated end date of at least November 21, 2026, after which the cost comparison in this report may no longer hold.


Sonnet 5's strongest published case is breadth: it is the only one of the three with same-vendor benchmark figures across SWE-bench Pro, OSWorld, BrowseComp, and GDPval-AA v2, all directly measured against its own predecessor and against Opus 4.8, giving a team more surface area to match against a specific workload even where a cross-vendor number does not exist. Its cancelled price increase means the $2/$10 rate used throughout this report should be treated as stable rather than temporary, unlike Sol's discount and Gemini's scheduled hike, which is itself a reason to weight it more heavily in a multi-year infrastructure decision even without a Terminal-Bench win to point to.


·····

FOLLOW US FOR MORE.

·····

DATA STUDIOS

·····

[datastudios.org]

bottom of page