SpaceXAI has released Grok 4.7 at the same starting API rates as Grok 4.6: $2 per million input tokens and $6 per million output tokens below 200,000 tokens of context. The company reports higher benchmark scores, although most of its direct comparisons use Grok 4.7 at xHigh effort and Grok 4.6 at High. Independent testing also suggests that 4.7 can generate substantially more tokens, so equal token rates may not mean equal cost per task.

The model arrived on September 21, 2026, after the early-September target dates Elon Musk had posted passed without a release. This article is for anyone deciding whether to route work to Grok 4.7: the published rates, the benchmark table as SpaceXAI presented it, the one independent measurement available, and where the model can be used. We have not run our own tests.

What SpaceXAI says changed

The launch post describes a new, larger base model trained with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours. SpaceXAI says Grok 4.7 verifies its own work more carefully, manages long context better, and was trained to understand the Grok Bot harness natively, which the company links to better conversational and general knowledge work. Grok Bot is SpaceXAI’s agent product; our Grok Bot guide covers what that harness does.

SpaceXAI also describes a new safeguard stack and calls Grok 4.7 its strongest model on refusals and jailbreak resistance, citing 62.4 percent on LatchBio’s biosafety benchmark and results on HackerBench, a cyber benchmark SpaceXAI built itself.

Rates and specifications

Grok 4.7Grok 4.6
Input, per million tokens, under 200K context$2.00$2.00
Output, per million tokens, under 200K context$6.00$6.00
Input and output above 200K context$4.00 and $12.00$4.00 and $12.00
Cached input$0.50$0.50
Context window500,000 tokens500,000 tokens
Fast variant$4.00 and $12.00, twice the output speedNot listed
Image inputYes, up to 20 MiBYes
API model idgrok-4.7grok-4.6

Source: SpaceXAI’s models documentation. The rates are per token. What a task costs depends on how many tokens the model spends on it, which is where the two models differ.

The benchmark table, as published

SpaceXAI’s main comparison sets Grok 4.7 against Grok 4.6, OpenAI’s GPT-5.6 Sol, and Anthropic’s Fable 5.1. The effort levels in the column heads are SpaceXAI’s own labels, and they are not matched: Grok 4.7 at xHigh, Grok 4.6 at High, the rivals at Max. Higher effort means more reasoning tokens per task, so the gaps below mix model quality with spend.

BenchmarkGrok 4.7 (xHigh)Grok 4.6 (High)GPT-5.6 Sol (Max)Fable 5.1 (Max)
Rates, input and output per million$2 and $6$2 and $6$4 and $20$10 and $50
CursorBench 4.0, software engineering46.3%40.4%41.7%51.8%
DeepSWE v1.1, software engineering71.0%*65.2%72.7%70.0%
EEBench, electrical engineering64.0%53.0%39.4%56.4%
AA Briefcase v1.1, multi-hour office work1,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional, clinical reasoning56.7%48.5%60.5%62.1%

*SpaceXAI marks this one Grok 4.7 score as measured at High effort.

A separate chart adds GDPval, a professional knowledge-work benchmark scored in Elo: Fable 5.1 (max) 1,735, Grok 4.7 (xhigh) 1,695, Grok 4.6 (high) 1,605, GPT-6 Astra (max) 1,542. That is the only chart in which GPT-6 Astra, OpenAI’s current flagship since September 3, 2026, appears; the main table uses the older GPT-5.6 Sol.

What independent testing shows

Artificial Analysis, which runs its own ten-benchmark Intelligence Index, had scored Grok 4.7 by September 22.

Measure (Artificial Analysis, index v4.3.2)Grok 4.7 HighGrok 4.7 xHighGrok 4.6 High
Intelligence Index464644
Output tokens generated across the evaluation suitenot listed240 million94 million
Suite median for comparison92 million92 million

On this index the effort setting makes no difference to Grok 4.7’s score; whether that holds on CursorBench or the legal benchmark is unknown, but it is the only matched-effort data point available. The xHigh run also generated about 2.5 times the output tokens of Grok 4.6 at High. If that carries into real workloads, the same per-token rate buys fewer completed tasks per dollar. Artificial Analysis scores GPT-6 Astra at 53 on the same index at roughly five times Grok’s token rates.

What the launch post leaves out

  • Consumer apps. Availability is Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers. Nothing is said about grok.com, the iOS and Android apps, or SuperGrok subscribers.
  • Pre-launch claims. Musk’s September posts made claims about the model’s size and training data that the launch post does not repeat, and SpaceXAI has published no parameter count or training details.
  • Vision. Musk said in mid-September that multimodal performance needed work. The launch post lists image input but no vision benchmark.
  • A model card or safety report. The post links to the API console and docs, not to a system card.

How to test it fairly

  • Already on Grok 4.6 through the API: run a fixed batch of your real tasks on both models at the same effort level, High against High, and compare quality alongside total tokens billed. The rate table alone will not tell you which is cheaper.
  • Using Cursor or Grok Build: Grok 4.7 is selectable now. Watch token consumption per task as closely as output quality.
  • Waiting on the consumer app: nothing has been announced. Our guide to using Grok covers the current app.

FAQ

Is Grok 4.7 available on grok.com or in the Grok app?
SpaceXAI’s launch post lists Cursor, Grok Build, the Grok API, third-party coding tools, and model routers. It does not mention grok.com, the mobile apps, or SuperGrok plans, and no rollout date for them has been announced.
What is the Grok 4.7 Fast variant?
SpaceXAI serves a second version of Grok 4.7 with twice the output speed at twice the rate: $4 per million input tokens and $12 per million output tokens, against $2 and $6 for the standard model.