Grok 4.6 vs Grok 4.7: Complete Comparison and Report on Pricing, Benchmarks, Independent Testing, and What Actually Changed

xAI released Grok 4.7 on September 21, 2026, forty days after Grok 4.6 and thirty-one days after the date Elon Musk had originally targeted for it, following four separate delays. The launch pitch was straightforward: a larger base model, longer reinforcement learning on harder multi-hour tasks, and pricing left exactly where Grok 4.6 set it.
What followed within hours was a split between what xAI's own launch table claims and what the first independent measurement found. Artificial Analysis scored Grok 4.7 at 46 on its Intelligence Index, ranked 16th of 655 evaluated models, a result that sits well below Grok 4.6's own score of 61 on the same index. This report sets xAI's claims against that independent reading and against everything else published in the first day after launch.
··········
RELEASE TIMELINE AND WHAT XAI CHANGED.
Model identifiers, lineage, and the basic facts of what shipped.
........
Attribute | Grok 4.6 | Grok 4.7 |
Release date | August 12, 2026 | September 21, 2026 |
Model ID | grok-4.6 | grok-4.7 |
Base model | Same base as Grok 4.5, post-training refinement only | New, larger base model per xAI |
Parameter count | Approximately 1.5 trillion | 2.1 trillion claimed by Elon Musk on X; not confirmed in xAI's official launch documentation |
Knowledge cutoff | February 1, 2026 | June 2026 |
Days since predecessor | 35 days (from Grok 4.5) | 40 days (from Grok 4.6) |
Access | xAI API, Grok Bot, Grok Build, Cursor | xAI API, Grok Build, Cursor, GitHub Copilot |
........
The parameter count is worth treating carefully. The 2.1 trillion figure, a stated 40% increase over Grok 4.6, comes from Musk's own social media posts on launch day, not from xAI's model card or launch post, which does not confirm a specific number. Grok 4.6 itself was explicitly built on the same base as Grok 4.5 rather than a larger one, so if the 2.1 trillion figure holds, Grok 4.7 would be the first genuine base-model scale-up in this line since 4.5, rather than the post-training-only refinement that characterized the 4.5-to-4.6 step.
··········
PRICING: IDENTICAL RATES, DIFFERENT REALITIES.
The headline number did not move. What it takes to complete a task did.
........
Rate | Grok 4.6 | Grok 4.7 |
Standard input (under 200,000 tokens) | $2 per million tokens | $2 per million tokens |
Standard output (under 200,000 tokens) | $6 per million tokens | $6 per million tokens |
Cached input | $0.50 per million tokens | $0.50 per million tokens |
Above 200,000-token threshold | $4 input, $12 output, $1 cached | $4 input, $12 output, $1 cached |
Fast variant | Not published as a distinct tier at 4.6's launch | 2x standard price, roughly 2x output speed |
US regional endpoint surcharge | Not published in this comparison | +10% on token prices |
........
The rate card is unchanged line for line. What changed is how many tokens a task consumes to reach an answer. Artificial Analysis's Intelligence Index run measured Grok 4.7 using 81,000 output tokens per task, against 36,000 for Grok 4.6 on the same evaluation, a 125% increase. At an unchanged per-token price, a task that costs the same on paper can cost more than double in practice if it now takes the model more than twice as many tokens to finish it.
Grok 4.7's Fast variant is also structured differently than any prior tier: it is available through Cursor and Grok Build but not through the public xAI API, and it is excluded from Grok Build's free entry tier. A developer calling the API directly cannot currently reach the faster, more expensive version at all.
··········
CONTEXT WINDOW, MODALITY, AND KNOWLEDGE CUTOFF.
What stayed fixed and what moved forward.
........
Specification | Grok 4.6 | Grok 4.7 |
Context window | 500,000 tokens | 500,000 tokens |
Input modalities | Text, images | Text, images |
Reasoning effort levels | Low, medium, high, xhigh | Not separately published; xAI markets longer reasoning by default |
Knowledge cutoff | February 1, 2026 | June 2026 |
........
The context window is identical and remains half the size of the roughly 1-million-token windows this outlet has documented on GPT-6 Astra, Claude Opus 5, and Claude Fable 5.1. The knowledge cutoff moved forward by four months, a routine update that reflects newer training data rather than an architectural change.
··········
XAI'S OWN BENCHMARK TABLE: WHAT THE VENDOR CLAIMS.
The comparison xAI chose to publish, including against two other vendors' models.
xAI's launch table reports Grok 4.7 beating Grok 4.6 on all seven benchmarks the company chose to publish, and beating GPT-5.6 Sol Max on four of those seven. Against Claude Fable 5.1 Max, the comparison runs both ways: Fable 5.1 Max leads Grok 4.7 on CursorBench, Terminal-Bench, GDPval, and HealthBench, while Grok 4.7 leads Fable on EEBench and the Harvey legal benchmark. xAI's own characterization of this mixed result is "highly competitive in its class."
On CursorBench 4.0 specifically, which evaluates long-duration coding tasks, xAI reports Grok 4.7 at 46.3%, a gain of 5.9 percentage points over the figure it reports for Grok 4.6. Separately, xAI states Grok 4.7 outperforms GPT-5.6 Sol on CursorBench 4.0 and edges out Claude Fable 5.1 Max on DeepSWE v1.1.
··········
THE ARTIFICIAL ANALYSIS INTELLIGENCE INDEX: AN INDEPENDENT REGRESSION.
The single clearest apples-to-apples measurement available so far, and it does not match the vendor's framing.
........
Model | Artificial Analysis Intelligence Index | Rank | Output tokens per task |
Grok 4.6 | 61 | Not published in this comparison | 36,000 |
Grok 4.7 | 46 | 16th of 655 models evaluated | 81,000 |
........
This is the one benchmark in this report where the same independent evaluator scored both models on the same methodology, which makes the 15-point gap the most reliable single data point available at the time of this report. A model using more than double the output tokens per task while scoring substantially lower on a composite reasoning-and-knowledge index is a result that runs directly against xAI's own claim of beating Grok 4.6 across the board.
This does not necessarily mean Grok 4.7 is a weaker model in every respect; xAI's benchmark selection and Artificial Analysis's index measure different things, and the sections below show real, independently plausible gains on specific coding benchmarks. But it does mean that "beats its predecessor on every published benchmark," xAI's own framing, does not extend to the one index an outside party has measured on both models under identical conditions.
··········
TERMINAL-BENCH VERSION CONFUSION ACROSS TWO RELEASES.
A recurring measurement problem that predates Grok 4.7 by one generation.
Terminal-Bench comparisons for this model line have carried a version mismatch since before Grok 4.7 existed. At Grok 4.6's own launch, xAI reported 26% on Terminal-Bench v3.0, a harder revision, while Artificial Analysis separately reported 88.39% for the same model on Terminal-Bench v2.1, an easier one. Both figures were accurate; they simply described different tests.
The same pattern recurs at Grok 4.7's launch. One outlet's coverage describes xAI's own claim as "nearly doubling" on Terminal-Bench, while Artificial Analysis's independently measured Terminal-Bench 4.0 score for Grok 4.7 is 26%, a figure separately reported as trailing GPT-6 Astra's roughly 60% and Claude Fable 5.1's roughly 55% by a wide margin. No source reviewed for this report published a Terminal-Bench 4.0 score for Grok 4.6 specifically, so the 26% figure for 4.7 cannot currently be checked against a same-version predecessor score at all. Readers should treat any Terminal-Bench comparison across these two models as unverifiable until both are scored on a stated, matching version.
··········
AGENTIC CODING BENCHMARKS: CURSORBENCH AND DEEPSWE.
Where Grok 4.7's gains are more directly supported.
........
Benchmark | Grok 4.6 | Grok 4.7 |
CursorBench 4.0 | 40.4% (implied by xAI's stated 5.9-point gain) | 46.3% |
DeepSWE v1.1 | 65.9% (xAI-reported; independent listing shows regression from 4.5) | Edges out Claude Fable 5.1 Max, per xAI (exact score not published in this comparison) |
........
CursorBench is the one benchmark in this report with a clear, consistent stated gain across the two generations, and it is a benchmark specifically built around long-duration coding tasks, the category xAI says it targeted with 4.7's training. DeepSWE comparisons carry the same caution flagged in this outlet's earlier coverage of Grok 4.6: xAI's own reported score for 4.6 was contradicted by an independent LiveBench listing showing a decline from Grok 4.5, so a claimed edge over Fable 5.1 Max on the same benchmark for 4.7 should be treated as a vendor claim pending independent confirmation.
··········
A RECURRING PATTERN: GROK 4.6'S OWN INDEPENDENT REGRESSION.
This is not the first time this model line has shown a gap between vendor claims and outside measurement.
At Grok 4.6's own launch, an early LiveBench listing reported results in the opposite direction from xAI's launch claims on two separate measures: agentic coding down to 54.2 from Grok 4.5's 56.5, and SkillsBench down to 55.77 from 66.03, alongside a time-to-first-token increase from 8.7 to 31.2 seconds. That dispute was never fully resolved by further independent testing in the material available at the time.
Placed side by side, Grok 4.6 and Grok 4.7 now each carry a documented instance of an independent source measuring a regression on a specific benchmark in the same release cycle where xAI's own materials claimed improvement. One instance is a data point; two consecutive releases showing the same pattern is closer to a house style, and it is a reasonable basis for treating any single-source vendor benchmark table from this company as provisional until a second party checks it.
··········
SAFETY STACK AND SECURITY BENCHMARKS.
New disclosures at 4.7, with no directly comparable 4.6 figures.
xAI describes Grok 4.7 as carrying an entirely new safeguard stack and states it is the company's strongest model yet on refusals and resistance to jailbreaks. Two specific figures accompany that claim: 3.3% of risky dual-use cyber prompts were allowed through on HackerBench v0.3, and the model scored 62.4% on LatchBio's biosafety benchmark.
No equivalent figures for Grok 4.6 on either benchmark were located for this report, which means neither number can currently be read as an improvement or a regression relative to the prior generation, only as a new baseline. This mirrors a disclosure gap this outlet has flagged in other safety-focused comparisons: a new safety claim without a same-benchmark predecessor score is a starting point for future tracking, not yet a measured trend.
··········
TOOL PRICING AND ACCESS CHANGES.
Costs that shifted around the model rather than in its base rate.
xAI raised the price of its X Search tool effective September 21, 2026 at noon Pacific time, the same day as the Grok 4.7 launch, splitting the rate into $5 per 1,000 posts fetched and $10 per 1,000 profiles fetched. This is a different structure from the flat per-call rate this outlet documented for Grok 4.6, and it means a search-heavy workload can now cost more per call than it did under the prior pricing, independent of which model generation is making the calls.
The new 10% surcharge on the US regional API endpoint is a separate addition with no stated equivalent at Grok 4.6's launch, and it applies specifically to the regional endpoint rather than to xAI's default routing.
··········
RELEASE TIMELINE SLIPPAGE AND WHAT IT SIGNALS.
Four delays against a date the founder set publicly.
Grok 4.7 shipped 40 days after Grok 4.6, a shorter interval than the 35 days between Grok 4.5 and 4.6, but only after missing Musk's original public target date four separate times. A pattern of repeated public delays followed by a release that maintains unchanged pricing and shows a measured regression on at least one independent index is worth weighing against the marketing framing that accompanied the launch, which called the model a combination of higher intelligence, speed, and low cost without qualification.
··········
WHAT ACTUALLY CHANGED: A FINAL ASSESSMENT.
Three things changed with confidence: the knowledge cutoff moved from February to June 2026, CursorBench 4.0 rose from a stated 40.4% to 46.3%, and the X Search tool got more expensive on the same day as the model launch. Everything else in this comparison carries a qualification. The parameter count increase to 2.1 trillion rests on a founder's social media post rather than a confirmed technical document. The Terminal-Bench figures for both generations describe different benchmark versions and cannot be compared to each other at all. The new safety benchmarks have no predecessor baseline to measure against. And the one clean, same-methodology, independently run comparison available, Artificial Analysis's Intelligence Index, shows Grok 4.7 scoring 15 points lower than Grok 4.6 while consuming more than double the output tokens per task.
None of this means Grok 4.7 is a worse model than Grok 4.6 in every practical sense; CursorBench's stated gain is real and specific to the long-duration coding work xAI says it targeted. It does mean that a decision to move from 4.6 to 4.7 should be based on testing the specific workload in question against both models directly, rather than on xAI's aggregate claim of a universal upgrade, given that the one outside party to test both under matching conditions found the opposite result on its own index.
·····
FOLLOW US FOR MORE.
·····
DATA STUDIOS
·····
[datastudios.org]




