SpaceXAI has released Grok 4.7 at the same starting API rates as Grok 4.6: $2 per million input tokens and $6 per million output tokens below 200,000 tokens of context. The company reports higher benchmark scores, although most of its direct comparisons use Grok 4.7 at xHigh effort and Grok 4.6 at High. Independent testing also suggests that 4.7 can generate substantially more tokens, so equal token rates may not mean equal cost per task.
The model arrived on September 21, 2026, after the early-September target dates Elon Musk had posted passed without a release. This article is for anyone deciding whether to route work to Grok 4.7: the published rates, the benchmark table as SpaceXAI presented it, the one independent measurement available, and where the model can be used. We have not run our own tests.
What SpaceXAI says changed
The launch post describes a new, larger base model trained with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours. SpaceXAI says Grok 4.7 verifies its own work more carefully, manages long context better, and was trained to understand the Grok Bot harness natively, which the company links to better conversational and general knowledge work. Grok Bot is SpaceXAI’s agent product; our Grok Bot guide covers what that harness does.
SpaceXAI also describes a new safeguard stack and calls Grok 4.7 its strongest model on refusals and jailbreak resistance, citing 62.4 percent on LatchBio’s biosafety benchmark and results on HackerBench, a cyber benchmark SpaceXAI built itself.
Rates and specifications
| Grok 4.7 | Grok 4.6 | |
|---|---|---|
| Input, per million tokens, under 200K context | $2.00 | $2.00 |
| Output, per million tokens, under 200K context | $6.00 | $6.00 |
| Input and output above 200K context | $4.00 and $12.00 | $4.00 and $12.00 |
| Cached input | $0.50 | $0.50 |
| Context window | 500,000 tokens | 500,000 tokens |
| Fast variant | $4.00 and $12.00, twice the output speed | Not listed |
| Image input | Yes, up to 20 MiB | Yes |
| API model id | grok-4.7 | grok-4.6 |
Source: SpaceXAI’s models documentation. The rates are per token. What a task costs depends on how many tokens the model spends on it, which is where the two models differ.
The benchmark table, as published
SpaceXAI’s main comparison sets Grok 4.7 against Grok 4.6, OpenAI’s GPT-5.6 Sol, and Anthropic’s Fable 5.1. The effort levels in the column heads are SpaceXAI’s own labels, and they are not matched: Grok 4.7 at xHigh, Grok 4.6 at High, the rivals at Max. Higher effort means more reasoning tokens per task, so the gaps below mix model quality with spend.
| Benchmark | Grok 4.7 (xHigh) | Grok 4.6 (High) | GPT-5.6 Sol (Max) | Fable 5.1 (Max) |
|---|---|---|---|---|
| Rates, input and output per million | $2 and $6 | $2 and $6 | $4 and $20 | $10 and $50 |
| CursorBench 4.0, software engineering | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1, software engineering | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench, electrical engineering | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1, multi-hour office work | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional, clinical reasoning | 56.7% | 48.5% | 60.5% | 62.1% |
*SpaceXAI marks this one Grok 4.7 score as measured at High effort.
A separate chart adds GDPval, a professional knowledge-work benchmark scored in Elo: Fable 5.1 (max) 1,735, Grok 4.7 (xhigh) 1,695, Grok 4.6 (high) 1,605, GPT-6 Astra (max) 1,542. That is the only chart in which GPT-6 Astra, OpenAI’s current flagship since September 3, 2026, appears; the main table uses the older GPT-5.6 Sol.
What independent testing shows
Artificial Analysis, which runs its own ten-benchmark Intelligence Index, had scored Grok 4.7 by September 22.
| Measure (Artificial Analysis, index v4.3.2) | Grok 4.7 High | Grok 4.7 xHigh | Grok 4.6 High |
|---|---|---|---|
| Intelligence Index | 46 | 46 | 44 |
| Output tokens generated across the evaluation suite | not listed | 240 million | 94 million |
| Suite median for comparison | 92 million | 92 million |
On this index the effort setting makes no difference to Grok 4.7’s score; whether that holds on CursorBench or the legal benchmark is unknown, but it is the only matched-effort data point available. The xHigh run also generated about 2.5 times the output tokens of Grok 4.6 at High. If that carries into real workloads, the same per-token rate buys fewer completed tasks per dollar. Artificial Analysis scores GPT-6 Astra at 53 on the same index at roughly five times Grok’s token rates.
What the launch post leaves out
- Consumer apps. Availability is Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers. Nothing is said about grok.com, the iOS and Android apps, or SuperGrok subscribers.
- Pre-launch claims. Musk’s September posts made claims about the model’s size and training data that the launch post does not repeat, and SpaceXAI has published no parameter count or training details.
- Vision. Musk said in mid-September that multimodal performance needed work. The launch post lists image input but no vision benchmark.
- A model card or safety report. The post links to the API console and docs, not to a system card.
How to test it fairly
- Already on Grok 4.6 through the API: run a fixed batch of your real tasks on both models at the same effort level, High against High, and compare quality alongside total tokens billed. The rate table alone will not tell you which is cheaper.
- Using Cursor or Grok Build: Grok 4.7 is selectable now. Watch token consumption per task as closely as output quality.
- Waiting on the consumer app: nothing has been announced. Our guide to using Grok covers the current app.