xAI announced Grok 4.7 on September 21, 2026. The list price did not move: $2 per million input tokens and $6 per million output, the same figures xAI published for Grok 4.6 in August. What changed is the model underneath, the numbers on xAI's own benchmark table, and a paid speed tier that is only sold through two products.
This post is for anyone choosing a coding model this week. It prints xAI's comparison table with a column that says who ran each test, checks the price a developer is billed on the routes most coding harnesses use, and ends with three questions that decide whether the switch is worth making. Every benchmark figure in the table below is xAI's own measurement and is labelled as such.
- 01Same list price as Grok 4.6, on the same tiers: $2/$6 below 200,000 prompt tokens, $4/$12 above.xAI's September release note states the tiers and the $0.50 cached-input rate. A fast variant runs at twice those rates and is sold only through Cursor and Grok Build, not on the public API.
- 02On xAI's table Grok 4.7 gains on all seven rows over 4.6 and trails Claude Fable 5.1 on the two coding rows.CursorBench 4.0: 46.3% against Fable 5.1's 51.8%. Terminal-Bench 4.0: 38.0% against 57.9%. It leads the field on EEBench and the Harvey legal benchmark. Vendor-run, no error bars published.
- 03The route price on OpenRouter was $1.60/$4.80 on September 22, 20% under the list price.The same 2x step above 200,000 prompt tokens applies on the route. Two priority endpoints charge double. GitHub Copilot bills the model at provider list pricing under usage-based billing.
- 04Switch a harness only if your own task set, at your own effort level, shows the gain.The asterisked DeepSWE figure is a high-effort run, and effort is the variable that moves both the score and the bill.
01 — The short versionWhat changed, in two lines
First line: a new, larger base model, trained with a longer reinforcement learning run weighted toward tasks that take hours. That is xAI's description in its announcement, and it is as specific as xAI gets about the architecture. No parameter count is published. The context window is 500,000 tokens, stated in the developer release notes, with text and image input and text-only output.
Second line: nothing changed on the invoice for API users, and one thing was added for two products. The release note prices the model at $2 input, $0.50 cached input and $6 output per million tokens below 200,000 prompt tokens, doubling to $4, $1 and $12 above that line. Those are the Grok 4.6 tiers to the cent. The addition is Grok 4.7 Fast, which xAI describes as the same model at twice the token rates and twice the output speed, available only in Cursor and Grok Build.
Reasoning effort has four levels, low, medium, high and xhigh, with high as the default. That matters for reading the table below, because xAI reports its headline coding scores at xhigh and one at high. Our effort-ladder reference covers what each vendor's levels mean and how they differ.
Served at the same price and speed as Grok 4.6, it is highly competitive in its class.xAI, Introducing Grok 4.7, September 21, 2026
02 — The numbersThe benchmark table, with its provenance
Every figure in this table was measured and published by xAI on September 21, 2026. That is the only source for any of them. The competitor scores are xAI's runs of the competitor models, not those vendors' own published numbers; where a vendor has published a different score for the same test, the difference is the harness and the run, and neither figure is independent.
The effort levels are not uniform. Grok 4.7 is reported at xhigh except on DeepSWE, where the asterisk in xAI's table marks a high-effort score. Grok 4.6 is at high, GPT-5.6 Sol at max and Fable 5.1 at max. xAI publishes no error bars and no run counts.
| Benchmark | Grok 4.7 | Grok 4.6 | Best other, per xAI |
|---|---|---|---|
| CursorBench 4.0 (software engineering) | 46.3% | 40.4% | 51.8% (Fable 5.1) |
| DeepSWE v1.1 (software engineering) | 71.0% at high effort | 65.2% | 72.7% (GPT-5.6 Sol) |
| EEBench (electrical engineering) | 64.0% | 53.0% | 56.4% (Fable 5.1) |
| AA Briefcase v1.1 (multi-hour office work) | 1,657 | 1,546 | 1,678 (Fable 5.1) |
| Terminal-Bench 4.0 (multi-hour terminal work) | 38.0% | 20.3% | 57.9% (Fable 5.1) |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 6.7% (Fable 5.1) |
| HealthBench Professional (clinical reasoning) | 56.7% | 48.5% | 62.1% (Fable 5.1) |
Read across the rows and the shape is consistent. Grok 4.7 beats Grok 4.6 everywhere, by 3.8 to 17.7 points on the percentage rows. It leads all three other models on EEBench and on the Harvey legal benchmark. It trails Fable 5.1 on the two coding rows that a harness buyer looks at first, CursorBench 4.0 and Terminal-Bench 4.0, and on the office-work and clinical rows; on DeepSWE v1.1 it leads Fable 5.1 and trails GPT-5.6 Sol. The claim xAI actually makes is about price-performance on CursorBench 4.0, and at $2 against $10 input that claim is about the denominator.
One comparison to avoid: Anthropic's own launch post for Fable 5.1 reports a Terminal-Bench 4.0 score of 55.8%, and xAI reports 57.9% for the same model. Both are vendor runs on different days in different harnesses. Neither corrects the other, and a table that mixes them is not a table.
xAI reports 62.4% on LatchBio's biosafety benchmark and says Grok 4.7 tops it. On HackerBench v0.3, which xAI describes as its own benchmark for risky and malicious cyber tasks, it reports that 3.3% of risky dual-use prompts were allowed through. Both are xAI-run, one on xAI's own test. We print them because the announcement leads with them; we make no claim beyond the two figures and their names.
03 — The invoiceThe price you are actually billed
A developer on a coding harness rarely pays a vendor's list price directly. The harness or a router sits in between, and the number on that route is the one that lands on the bill. We read the OpenRouter route for Grok 4.7 on September 22, 2026, the day after launch. The values below are that observation, dated to the reading; a route can change without notice.
- xAI list price, below 200,000 prompt tokensInput / cached input / output per million tokens. Source: xAI developer release notes, September 2026.
- $2 / $0.50 / $6
- xAI list price, at or above 200,000 prompt tokensThe same 2x step Grok 4.6 carried. The whole request is billed at the higher tier once the prompt crosses the line.
- $4 / $1 / $12
- OpenRouter base route, read September 22Input / cached input / output. 20% under the list on every line. Context 500,000, max output 450,000, reasoning mandatory, four effort levels.
- $1.60 / $0.40 / $4.80
- OpenRouter route above 200,000 prompt tokensThe route carries the same threshold as the list, at the same 2x multiplier.
- $3.20 / $0.80 / $9.60
- OpenRouter priority endpointsTwo of the four endpoints on the route are tagged priority and charge double the base. They are not labelled as the fast variant, and we do not assume they are.
- $3.20 / $9.60
- GitHub CopilotGitHub's changelog of September 21 says "This model is billed at provider list pricing under usage-based billing." That is xAI's $2/$6, not the route's $1.60/$4.80.
- List price
The practical reading: for agentic coding, the 200,000-token line is the number to watch, not the headline rate. A long session that crosses it is billed at $4 and $12 on the API, and the cached-input rate doubles with it. Our harness-cost post showed the same model costing up to five times as much depending on the harness around it, and a long-context surcharge is one of the mechanisms. The list and route prices for every current model are in our price index, which will carry the Grok 4.7 row with its check date.
04 — The fine printWhat "same price" hides
"Same price as Grok 4.6" is true of the base list rate. Three things sit outside it, and each one changes the answer for a particular kind of buyer. The third comes from GitHub's changelog entry of September 21.
The fast variant
xAI's release note says Grok 4.7 Fast is the same model at twice the token rates, and that it is not on the public xAI API. On xAI's models page that is $4 / $1 / $12 below 200,000 prompt tokens, twice the standard band, and $6 / $1.50 / $18 above it, which is 1.5 times the standard long-context band rather than double. There is no separate public route to read a price from, so "twice the price" is xAI's own statement and nothing more. If your harness offers a fast toggle, that toggle is a 2x bill.
The 200,000-token tier
At or above 200,000 prompt tokens every rate doubles. A coding agent that reads a large repository into context crosses that line early in a session and stays there. The list price a buyer compares is the one below the line.
Copilot's billing basis
GitHub added Grok 4.7 to Copilot on September 21 in VS Code, Visual Studio, the CLI, the cloud agent, the Copilot app, JetBrains, Xcode and Eclipse, rolling out gradually. It is billed at provider list pricing under usage-based billing, and enterprise and business administrators control access through the model policy.
05 — The decisionThree questions before you switch
A new row on a vendor's table is not a reason to change the model behind a team's coding harness. These three questions are, and each one has an answer only your own usage can give. We cover how to read pre-release evidence more generally in our post on launch signals.
If the answers favour the switch, make it on one team for two weeks with the previous model still configured as a fallback, and keep the eval set. That is the same practice we use when we set up model selection for clients, and it is what turns a vendor's table into a decision you can defend.
06 — ConclusionA better model at the old price is still a vendor's claim until your tasks confirm it
Run your own task set at both effort levels, check the route price on the day, and decide from those two numbers
Grok 4.7 is a real step over Grok 4.6 on every row of xAI's table, at an unchanged list price, and it is behind Fable 5.1 on the coding rows most buyers weigh. The fast variant, the 200,000-token tier and Copilot's list-price billing are where the bill moves. None of that is a reason to switch or to stay; your own eval set and your own usage shape are.