Claude Sonnet 5 vs GPT-5.6 Sol vs Gemini 3.8 Flash: Complete Comparison and Report on Pricing, Benchmarks, Context Window, and Value-Tier Positioning

Three models currently anchor the value tier at the three largest labs: Claude Sonnet 5, launched by Anthropic on June 30, 2026; GPT-5.6 Sol, launched by OpenAI on July 9, 2026; and Gemini 3.8 Flash, launched by Google on September 2, 2026. Unlike a flagship-versus-flagship comparison, these three were built to compete directly on cost per completed task, and each vendor explicitly frames its release around getting near-frontier results at a fraction of frontier pricing.
This report also has to flag something upfront: pricing on two of these three models has moved since launch, in opposite directions, and getting the current rate wrong changes the entire comparison.
··········
RELEASE TIMELINE AND MODEL IDENTITY.
Model identifiers, lineage, and how each vendor positions its value tier.
........
Attribute | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
Vendor | Anthropic | OpenAI | Google DeepMind |
Release date | June 30, 2026 | July 9, 2026 (limited preview June 26) | September 2, 2026 |
Model ID | claude-sonnet-5 | gpt-5.6-sol | gemini-3.8-flash |
Predecessor | Claude Sonnet 4.6 | GPT-5.5 | Gemini 3.7 Flash |
Position in lineup | Mid-tier, between Haiku and Opus | Flagship of the GPT-5.6 family (Sol/Terra/Luna) | Mid-tier, below the Pro line |
Base model status | New generation | New generation, distilled with Terra and Luna from one training run | Built on the Gemini 3.7 Flash architecture |
........
Only Sol carries flagship status inside its own family; Sonnet 5 and Gemini 3.8 Flash are both deliberately positioned as the economical option below a more expensive sibling tier at their respective labs. That distinction matters for reading the benchmark tables later in this report: a value-tier model losing to a frontier flagship in another vendor's lineup is not surprising, but a value-tier model beating a same-week flagship on a shared benchmark, which happens more than once below, is worth noting explicitly.
··········
PRICING AND THE TWO RATE CHANGES SINCE LAUNCH.
Current confirmed rates, and the corrections needed on top of each launch announcement.
........
Rate per 1M tokens | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
Standard input (current) | $2 | $5 | $0.75 (through Dec 31, 2026) |
Standard output (current) | $10 | $30 | $3.75 (through Dec 31, 2026) |
Cached input | $0.20 | $0.50 | Not published in this table |
Promotional adjustment | Planned increase to $3/$15 on September 1, 2026, since cancelled | Cut by more than 20% from August 21, 2026, through at least November 21, 2026 | Doubles to $1.50/$7.50 on January 1, 2027 |
........
Sonnet 5 launched with a stated introductory period ending August 31, 2026, after which the rate was due to rise to $3 input and $15 output. Anthropic's own pricing page, rechecked independently on August 20 and again on August 26, still showed $2/$10 with no introductory label and no announced increase, and the step-up has since been described as cancelled. Anyone quoting Sonnet 5 at $3/$15 today is citing a change that did not happen.
Sol moved in the opposite direction: OpenAI cut its listed $5/$30 rate by more than 20% starting August 21, 2026, confirmed available through at least November 21. The exact discounted figure varies slightly across sources reporting the cut, but the direction and the 20%-plus magnitude are consistent across OpenAI's own pricing update.
Gemini 3.8 Flash is the only one of the three with a scheduled future increase rather than a past correction: its entire rate card doubles on January 1, 2027, a date that applies across Google's whole Flash 3.x family rather than to this release specifically.
At current confirmed rates, Gemini 3.8 Flash costs roughly a third of Sonnet 5's rate and about a fifteenth of Sol's discounted rate on input, with a comparable multiple on output.
··········
CONTEXT WINDOW, OUTPUT LIMITS, AND TOKENIZER EFFECTS.
Window sizes and a cost variable that does not show up in the headline price.
........
Limit | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
Context window | 1,000,000 tokens | 1,050,000 tokens | 1,048,576 tokens |
Maximum output | 128,000 tokens | 128,000 tokens | 65,536 tokens |
Knowledge cutoff | Not separately published for Sonnet 5 in this comparison | February 16, 2026 | March 2026 for most domains, January 2025 for some |
Long-context billing threshold | None | Above 272,000 input tokens: 2x input, 1.5x output, 1.25x cache write | None |
........
Gemini 3.8 Flash's output ceiling is exactly half of the other two, a constraint documented in Google's own model card for agent work that reasons extensively before producing final output. Sol carries the same 272,000-token long-context penalty structure this outlet has documented on its flagship sibling Astra, applied here to the value tier as well.
Sonnet 5 introduced an updated tokenizer at launch, the same generation Anthropic later used for Opus 4.7 and Fable 5.1, which maps the same text to roughly 1.0 to 1.35 times more tokens depending on content type than Sonnet 4.6's tokenizer produced. Anthropic states it set the introductory price specifically to make the move from Sonnet 4.6 roughly cost-neutral despite the token-count increase, which means a raw per-token price comparison against Sonnet 4.6, or against any model using an older tokenizer, understates Sonnet 5's real cost per unit of text.
··········
BENCHMARK METHODOLOGY AND SOURCE RELIABILITY.
Why several of the most-cited figures for these three models need a version or date check before use.
None of the three vendors tested its value-tier model against both of the other two under matching conditions at launch. Sonnet 5's system card compares it against GPT-5.5 and Gemini 3.5 Flash, since GPT-5.6 Sol had not been released as of June 30, 2026; any Sonnet-5-versus-Sol comparison circulating today is a later, unofficial pairing rather than a vendor-published one.
Sol's own benchmark reporting is not fully consistent across sources checked for this report. One tracker lists Sol's Terminal-Bench 2.1 score at 88.80%, ranked first among tracked models on that specific benchmark, while a separate outlet's chart compares Sol against unnamed competitors without a table, and a third source's launch-day coverage frames a different model's Terminal-Bench 2.1 figure (83.4% for GPT-5.5) in a way that could be mistaken for Sol's own score if read quickly. The 88.80% figure is treated as Sol's in the table below because it is sourced to a benchmark-specific tracker rather than a comparative launch chart.
Gemini 3.8 Flash's DeepSWE score carries the same version ambiguity documented in this outlet's earlier Astra-versus-Gemini report: the model card cites version 1 of the benchmark, while Google's launch post cites version 1.1, treating the two as equivalent.
··········
DIRECTLY COMPARABLE SCORES.
The one benchmark where all three models have a same-version, independently trackable figure.
........
Benchmark | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
Terminal-Bench 2.1 | 80.4% (beats Opus 4.8's 74.6% on this specific benchmark) | 88.80% | 90.8% (up from 81.6% for Gemini 3.7 Flash) |
SWE-bench Pro | 63.2% | Not published in this comparison | Not published in this comparison |
OSWorld-Verified / OSWorld 2.0 | 81.2% | Not published in this comparison | Not published in this comparison |
BrowseComp (agentic search) | 84.7% | Not published in this comparison | Not published in this comparison |
GDPval-AA v2 (Elo) | 1,618 | Not published in this comparison | Not published in this comparison |
FrontierMath v2 | Not published in this comparison | 89.12% | Not published in this comparison |
Context Arena | Not published in this comparison | 97.63% | Not published in this comparison |
........
Terminal-Bench 2.1 is the only row with a plausible same-version reading across all three vendors, and on it, the two cheapest models in this comparison outscore the most expensive one: Gemini 3.8 Flash at 90.8% and Sol at 88.80% both beat Sonnet 5's 80.4%, despite Sonnet 5 costing more per token than Gemini and, until August, roughly the same as Sol's original list price. Every other row in this table has data from only one vendor, which means no ranking beyond this single benchmark can currently be supported across all three models.
··········
LATENCY AND THROUGHPUT.
Independently measured speed where available, against undisclosed figures where it is not.
........
Measure | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
p95 time to first token | Not published in this comparison | 8.37 seconds | 13.30 seconds |
Output throughput rank | Not published in this comparison | Not published in this comparison | 3rd of 196 models evaluated |
Median TTFT across evaluated models | Not published in this comparison | Not published in this comparison | 2.99 seconds |
........
Sol's 8.37-second time to first token sits closer to the cross-model median than Gemini 3.8 Flash's 13.30 seconds, though the two figures come from different measurement sources and are not guaranteed to use identical methodology. No public latency figure was located for Sonnet 5 in this research pass, which is itself worth noting given how central latency is to a model marketed for high-volume, everyday use.
··········
AVAILABILITY ACROSS PLATFORMS.
Where each model actually runs today.
........
Channel | Claude Sonnet 5 | GPT-5.6 Sol | Gemini 3.8 Flash |
Default consumer tier | Free and Pro plans on claude.ai | Plus, Pro, Business, Enterprise in ChatGPT | Not specified as a ChatGPT/Claude-style consumer default |
Coding agent | Claude Code | Codex | Not listed |
Third-party IDE integration | Cursor, VS Code, GitHub Copilot | Not specified | Not specified |
Cloud platforms | Claude Platform, Amazon Bedrock, Google Vertex AI, Microsoft Foundry | OpenAI API, Amazon Bedrock, Azure | Google Cloud |
Enterprise tiers | Max, Team, Enterprise | Business, Enterprise | Not specified in this comparison |
........
Sonnet 5 has the broadest documented third-party IDE presence of the three, appearing natively in Cursor, VS Code, and GitHub Copilot at launch. Sol and Gemini 3.8 Flash are each available on their vendor's own cloud plus at least one additional major cloud provider, while Sonnet 5 is the only one of the three confirmed on all three major clouds.
··········
WORKLOAD ALLOCATION CRITERIA.
What the corrected pricing and the single comparable benchmark actually support as a decision.
The pricing corrections in this report matter more than any single benchmark gap: at current rates, Gemini 3.8 Flash costs roughly a third of Sonnet 5 and a fifteenth of discounted Sol on input tokens, and on the one benchmark where all three can be read against the same version, it does not trail either competitor. That combination makes it the reasonable default to test first for high-volume, cost-sensitive agent work, provided a team accounts for its 65,536-token output ceiling and confirms the January 2027 price doubling does not break the economics of a long-term commitment made today.
Sol's case rests on the benchmarks it publishes and the other two do not: FrontierMath v2 and Context Arena figures with no equivalent Sonnet 5 or Gemini 3.8 Flash score to compare against, plus its 272,000-token pricing threshold, which needs checking against actual prompt-length distribution before it changes the cost picture. Its post-discount price is the one to budget against, not its original $5/$30 list rate, and that discount carries a stated end date of at least November 21, 2026, after which the cost comparison in this report may no longer hold.
Sonnet 5's strongest published case is breadth: it is the only one of the three with same-vendor benchmark figures across SWE-bench Pro, OSWorld, BrowseComp, and GDPval-AA v2, all directly measured against its own predecessor and against Opus 4.8, giving a team more surface area to match against a specific workload even where a cross-vendor number does not exist. Its cancelled price increase means the $2/$10 rate used throughout this report should be treated as stable rather than temporary, unlike Sol's discount and Gemini's scheduled hike, which is itself a reason to weight it more heavily in a multi-year infrastructure decision even without a Terminal-Bench win to point to.
·····
FOLLOW US FOR MORE.
·····
DATA STUDIOS
·····
[datastudios.org]




