AI DevelopmentModel release7 min readPublished September 21, 2026

$2 / $6 unchanged · 7 vendor-run benchmark rows · 1 paid speed tier · the route price is lower than the list

Grok 4.7 Costs the Same as 4.6. What Actually Changed

xAI released Grok 4.7 on September 21, 2026 at Grok 4.6's $2/$6 list price. The vendor-run benchmark table, the route price you are billed, three questions.

DA
Digital Applied Team
Research and practical guidance
AnnouncedSeptember 21, 2026
Route price checkedSeptember 22, 2026

xAI announced Grok 4.7 on September 21, 2026. The list price did not move: $2 per million input tokens and $6 per million output, the same figures xAI published for Grok 4.6 in August. What changed is the model underneath, the numbers on xAI's own benchmark table, and a paid speed tier that is only sold through two products.

This post is for anyone choosing a coding model this week. It prints xAI's comparison table with a column that says who ran each test, checks the price a developer is billed on the routes most coding harnesses use, and ends with three questions that decide whether the switch is worth making. Every benchmark figure in the table below is xAI's own measurement and is labelled as such.

Key takeaways
  1. 01
    Same list price as Grok 4.6, on the same tiers: $2/$6 below 200,000 prompt tokens, $4/$12 above.xAI's September release note states the tiers and the $0.50 cached-input rate. A fast variant runs at twice those rates and is sold only through Cursor and Grok Build, not on the public API.
  2. 02
    On xAI's table Grok 4.7 gains on all seven rows over 4.6 and trails Claude Fable 5.1 on the two coding rows.CursorBench 4.0: 46.3% against Fable 5.1's 51.8%. Terminal-Bench 4.0: 38.0% against 57.9%. It leads the field on EEBench and the Harvey legal benchmark. Vendor-run, no error bars published.
  3. 03
    The route price on OpenRouter was $1.60/$4.80 on September 22, 20% under the list price.The same 2x step above 200,000 prompt tokens applies on the route. Two priority endpoints charge double. GitHub Copilot bills the model at provider list pricing under usage-based billing.
  4. 04
    Switch a harness only if your own task set, at your own effort level, shows the gain.The asterisked DeepSWE figure is a high-effort run, and effort is the variable that moves both the score and the bill.

01The short versionWhat changed, in two lines

First line: a new, larger base model, trained with a longer reinforcement learning run weighted toward tasks that take hours. That is xAI's description in its announcement, and it is as specific as xAI gets about the architecture. No parameter count is published. The context window is 500,000 tokens, stated in the developer release notes, with text and image input and text-only output.

Second line: nothing changed on the invoice for API users, and one thing was added for two products. The release note prices the model at $2 input, $0.50 cached input and $6 output per million tokens below 200,000 prompt tokens, doubling to $4, $1 and $12 above that line. Those are the Grok 4.6 tiers to the cent. The addition is Grok 4.7 Fast, which xAI describes as the same model at twice the token rates and twice the output speed, available only in Cursor and Grok Build.

Reasoning effort has four levels, low, medium, high and xhigh, with high as the default. That matters for reading the table below, because xAI reports its headline coding scores at xhigh and one at high. Our effort-ladder reference covers what each vendor's levels mean and how they differ.

Served at the same price and speed as Grok 4.6, it is highly competitive in its class.xAI, Introducing Grok 4.7, September 21, 2026

02The numbersThe benchmark table, with its provenance

Every figure in this table was measured and published by xAI on September 21, 2026. That is the only source for any of them. The competitor scores are xAI's runs of the competitor models, not those vendors' own published numbers; where a vendor has published a different score for the same test, the difference is the harness and the run, and neither figure is independent.

The effort levels are not uniform. Grok 4.7 is reported at xhigh except on DeepSWE, where the asterisk in xAI's table marks a high-effort score. Grok 4.6 is at high, GPT-5.6 Sol at max and Fable 5.1 at max. xAI publishes no error bars and no run counts.

Source: xAI, Introducing Grok 4.7, September 21, 2026. All rows are xAI-run. The last column shows the higher of xAI's two competitor figures, GPT-5.6 Sol and Claude Fable 5.1.
BenchmarkGrok 4.7Grok 4.6Best other, per xAI
CursorBench 4.0 (software engineering)46.3%40.4%51.8% (Fable 5.1)
DeepSWE v1.1 (software engineering)71.0% at high effort65.2%72.7% (GPT-5.6 Sol)
EEBench (electrical engineering)64.0%53.0%56.4% (Fable 5.1)
AA Briefcase v1.1 (multi-hour office work)1,6571,5461,678 (Fable 5.1)
Terminal-Bench 4.0 (multi-hour terminal work)38.0%20.3%57.9% (Fable 5.1)
Harvey Legal Agent Benchmark19.6%15.8%6.7% (Fable 5.1)
HealthBench Professional (clinical reasoning)56.7%48.5%62.1% (Fable 5.1)

Read across the rows and the shape is consistent. Grok 4.7 beats Grok 4.6 everywhere, by 3.8 to 17.7 points on the percentage rows. It leads all three other models on EEBench and on the Harvey legal benchmark. It trails Fable 5.1 on the two coding rows that a harness buyer looks at first, CursorBench 4.0 and Terminal-Bench 4.0, and on the office-work and clinical rows; on DeepSWE v1.1 it leads Fable 5.1 and trails GPT-5.6 Sol. The claim xAI actually makes is about price-performance on CursorBench 4.0, and at $2 against $10 input that claim is about the denominator.

One comparison to avoid: Anthropic's own launch post for Fable 5.1 reports a Terminal-Bench 4.0 score of 55.8%, and xAI reports 57.9% for the same model. Both are vendor runs on different days in different harnesses. Neither corrects the other, and a table that mixes them is not a table.

The two safety numbers, as xAI states them

xAI reports 62.4% on LatchBio's biosafety benchmark and says Grok 4.7 tops it. On HackerBench v0.3, which xAI describes as its own benchmark for risky and malicious cyber tasks, it reports that 3.3% of risky dual-use prompts were allowed through. Both are xAI-run, one on xAI's own test. We print them because the announcement leads with them; we make no claim beyond the two figures and their names.

03The invoiceThe price you are actually billed

A developer on a coding harness rarely pays a vendor's list price directly. The harness or a router sits in between, and the number on that route is the one that lands on the bill. We read the OpenRouter route for Grok 4.7 on September 22, 2026, the day after launch. The values below are that observation, dated to the reading; a route can change without notice.

xAI list price, below 200,000 prompt tokensInput / cached input / output per million tokens. Source: xAI developer release notes, September 2026.
$2 / $0.50 / $6
xAI list price, at or above 200,000 prompt tokensThe same 2x step Grok 4.6 carried. The whole request is billed at the higher tier once the prompt crosses the line.
$4 / $1 / $12
OpenRouter base route, read September 22Input / cached input / output. 20% under the list on every line. Context 500,000, max output 450,000, reasoning mandatory, four effort levels.
$1.60 / $0.40 / $4.80
OpenRouter route above 200,000 prompt tokensThe route carries the same threshold as the list, at the same 2x multiplier.
$3.20 / $0.80 / $9.60
OpenRouter priority endpointsTwo of the four endpoints on the route are tagged priority and charge double the base. They are not labelled as the fast variant, and we do not assume they are.
$3.20 / $9.60
GitHub CopilotGitHub's changelog of September 21 says "This model is billed at provider list pricing under usage-based billing." That is xAI's $2/$6, not the route's $1.60/$4.80.
List price

The practical reading: for agentic coding, the 200,000-token line is the number to watch, not the headline rate. A long session that crosses it is billed at $4 and $12 on the API, and the cached-input rate doubles with it. Our harness-cost post showed the same model costing up to five times as much depending on the harness around it, and a long-context surcharge is one of the mechanisms. The list and route prices for every current model are in our price index, which will carry the Grok 4.7 row with its check date.

04The fine printWhat "same price" hides

"Same price as Grok 4.6" is true of the base list rate. Three things sit outside it, and each one changes the answer for a particular kind of buyer. The third comes from GitHub's changelog entry of September 21.

1
The fast variant
Cursor and Grok Build only

xAI's release note says Grok 4.7 Fast is the same model at twice the token rates, and that it is not on the public xAI API. On xAI's models page that is $4 / $1 / $12 below 200,000 prompt tokens, twice the standard band, and $6 / $1.50 / $18 above it, which is 1.5 times the standard long-context band rather than double. There is no separate public route to read a price from, so "twice the price" is xAI's own statement and nothing more. If your harness offers a fast toggle, that toggle is a 2x bill.

Vendor-stated
2
The 200,000-token tier
API and route

At or above 200,000 prompt tokens every rate doubles. A coding agent that reads a large repository into context crosses that line early in a session and stays there. The list price a buyer compares is the one below the line.

Where agentic bills grow
3
Copilot's billing basis
Pro, Pro+, Max, Business, Enterprise

GitHub added Grok 4.7 to Copilot on September 21 in VS Code, Visual Studio, the CLI, the cloud agent, the Copilot app, JetBrains, Xcode and Eclipse, rolling out gradually. It is billed at provider list pricing under usage-based billing, and enterprise and business administrators control access through the model policy.

List, not route

05The decisionThree questions before you switch

A new row on a vendor's table is not a reason to change the model behind a team's coding harness. These three questions are, and each one has an answer only your own usage can give. We cover how to read pre-release evidence more generally in our post on launch signals.

Does the gain hold on your tasks at the effort you actually run?
Re-run your own 20 to 50 representative tasks at the default high effort and at xhigh, with the current model as the control. xAI's coding scores are at xhigh; a harness that defaults to high is running a different model in practice.
Your eval set
What share of your sessions cross 200,000 prompt tokens?
Pull last month's usage. If a meaningful share of requests run long, price the switch at $4/$12, not $2/$6, and compare against the incumbent's long-context rate rather than its headline.
Your billing export
Which price will you actually pay: list, route or a harness's own rate?
Read the route on the day you switch and record it. On September 22 the route was 20% under the list and Copilot was at list. Those three numbers diverge, and the one in your contract is the one that matters.
Your harness invoice

If the answers favour the switch, make it on one team for two weeks with the previous model still configured as a fallback, and keep the eval set. That is the same practice we use when we set up model selection for clients, and it is what turns a vendor's table into a decision you can defend.

06ConclusionA better model at the old price is still a vendor's claim until your tasks confirm it

What to do this week

Run your own task set at both effort levels, check the route price on the day, and decide from those two numbers

Grok 4.7 is a real step over Grok 4.6 on every row of xAI's table, at an unchanged list price, and it is behind Fable 5.1 on the coding rows most buyers weigh. The fast variant, the 200,000-token tier and Copilot's list-price billing are where the bill moves. None of that is a reason to switch or to stay; your own eval set and your own usage shape are.

Digital Applied

Choose coding models on your own evidence.

We build the eval sets, the cost models and the fallback configuration that let a team switch models when the numbers say so, not when a launch post does.

Model evaluationCost modellingHarness configuration
Your next project

A model decision you can defend

  • A task set that reflects your work
  • A price model that includes long context
  • A switch with a fallback
Questions and answers

Applying this post

No. xAI's release notes give the same tiers for both: $2 input, $0.50 cached input and $6 output per million tokens below 200,000 prompt tokens, and $4, $1 and $12 above. The route price on OpenRouter was 20% under the list on September 22, 2026, which is a route decision, not a price cut.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Grok Build Uploaded Entire Repos: What Actually Fixed It

Grok Build reportedly uploaded full repos and git history to xAI's cloud. The /privacy toggle never stopped it — a server-side flag did. Agent trust is infra.

July 14, 2026 · 12 minRead
AI Development

Grok 4.1: xAI Emotional AI Complete Guide

Master Grok 4.1: EQ-Bench #1 ranking, 65% hallucination reduction, Fast API access, xAI benchmarks, and comparison with GPT-5.2 and Claude Opus 4.5.

December 17, 2025 · 13 minRead
AI Development

Six Public Signals That an AI Model Launch Is Hours Away

Six signals that a model is shipping, five checkable and one not, from stealth tests to alias flips, with the API field for each and three launches scored.

September 21, 2026 · 9 minRead
AI Development

What "Open Source" Actually Includes for an AI Model

16 models in the open-model conversation, checked against the licence file, OSI's list and what else was released: weights, data, recipes, evaluation harness.

September 21, 2026 · 7 minRead
AI Development

AI-Built Forms: Keep User Input When Submission Fails

Test AI-built forms beyond a successful submit. Preserve valid input, explain errors and distinguish a rejected request from an outcome still unknown.

September 6, 2026 · 4 minRead
AI Development

Small AI-Built Tools: Set the Boundary Before You Build

Scope a small AI-built utility around clear inputs, outputs and limits. Decide what it should own, reject and preserve before it grows into a system.

September 6, 2026 · 4 minRead