Grok 4.7 Lands at Grok 4.6 Prices, With Higher Scores
SpaceXAI released Grok 4.7 on 21 September 2026 and calls it its most capable model for coding and knowledge work. The model is available today in Cursor and Grok Build, and it is served at the same price and speed as Grok 4.6: $2 per million input tokens and $6 per million output tokens. A faster variant costs double.
The headline on SpaceXAI's Grok 4.7 announcement reads "Twice as fast, at half the price of comparable models." That framing, not the benchmark table, is the part growth leaders should read twice.
What SpaceXAI announced on 21 September 2026
Grok 4.7 is a frontier model from SpaceXAI built for coding and knowledge work, sold through an API at a fixed price per million tokens. SpaceXAI says it works longer on difficult tasks, checks its own work more carefully, and ships with the company's best-calibrated safeguards to date.
Availability on day one covers Cursor and Grok Build, plus the Grok API, third-party coding harnesses, and model routers and cloud platforms. Grok Build can be tried for free.
Pricing starts at $2 per million input tokens and $6 per million output tokens. SpaceXAI also serves a fast variant with twice the output speed at twice the price. Red-team capabilities are invite-only, limited to select cybersecurity partners.
What is actually new inside Grok 4.7?
A larger base model
Grok 4.7 uses a new, larger base model compared with Grok 4.6. That is a change in the foundation, not a tuning pass on the previous release.
A longer reinforcement learning run
SpaceXAI trained Grok 4.7 with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The company says the result is a model better at verifying its own work and managing longer context.
Native understanding of the Grok Bot harness
Grok 4.7 was trained to natively understand the Grok Bot harness, which SpaceXAI says makes it better at conversational tasks and general knowledge work. For anyone tracking brand mentions inside chat surfaces, that line matters more than the coding scores. A model tuned for the conversational harness is a model tuned for the place consumers ask questions.
How does Grok 4.7 pricing compare with GPT-5.6 Sol Max and Fable 5.1 Max?
The table below uses the figures SpaceXAI published for Grok 4.7 xHigh on 21 September 2026. Input and output token prices are identical to Grok 4.6 High.
| Measure | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input price, $ per million tokens | $2 | $2 | $4 | $10 |
| Output price, $ per million tokens | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high effort) | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
Token prices and benchmark scores as published by SpaceXAI on 21 September 2026. Output tokens cost $6 per million on Grok 4.7, against $50 on Fable 5.1 Max.
On professional knowledge work measured by GDPval Elo, Grok 4.7 (xhigh) reaches 1695. That sits behind Fable 5.1 (max) at 1735, and ahead of Grok 4.6 (high) at 1605 and GPT-6 Astra (max) at 1542.
Where do rival models still score higher?
Grok 4.7 is not the top score on most of the tables SpaceXAI published. Fable 5.1 Max beats it on CursorBench 4.0 (51.8% against 46.3%), Terminal-Bench 4.0 (57.9% against 38.0%), AA Briefcase v1.1 (1,678 against 1,657) and HealthBench Professional (62.1% against 56.7%). GPT-5.6 Sol Max leads on DeepSWE v1.1 (72.7% against 71.0%) and HealthBench Professional (60.5%).
Where Grok 4.7 does lead, the margin is wide. On EEBench it scores 64.0% against 39.4% for GPT-5.6 Sol Max. On the Harvey Legal Agent Benchmark it scores 19.6% against 2.5% for GPT-5.6 Sol Max and 6.7% for Fable 5.1 Max.
Read against the output price, the pattern is clear. Grok 4.7 is positioned on price-performance, not on winning the leaderboard.
What flat frontier pricing changes for growth and search teams
Cost per prompt, not model quality, has been the binding constraint on most AI measurement programmes I have seen. When a frontier model holds at $2 input and $6 output per million tokens, the maths behind several standing decisions changes.
- AI visibility sampling frequency: if you track brand mentions and citations across a few thousand prompts, output tokens dominate your bill. At $6 per million output tokens against $50 on Fable 5.1 Max, weekly or daily sampling becomes affordable where monthly was the ceiling.
- Bulk content QA: checking titles, schema, internal links and factual claims across a large site is a long, repetitive job. SpaceXAI says Grok 4.7 was trained on tasks weighted toward problems that take many hours, and is better at verifying its own work. That is the exact shape of a QA queue.
- Agentic site audits: Terminal-Bench 4.0 rose from 20.3% on Grok 4.6 to 38.0% on Grok 4.7, which is the closest public proxy for multi-hour tool-using work. It is still well behind Fable 5.1 Max at 57.9%, so treat crawl and log-file agents as assisted, not autonomous.
- Document and deck production: SpaceXAI says Grok 4.7 is better at creating documents and presentations, scoring 1,657 on AA Briefcase v1.1 against 1,546 for Grok 4.6. Reporting and board-pack drafting is the cheapest place to test it.
The second-order effect is the harness. Grok 4.7 was trained to natively understand the Grok Bot harness, and conversational surfaces are where a lot of brand discovery now happens. Most SEO teams have no measurement at all on chat assistants. My read is that the gap between what you rank for and what you get quoted in is now the more expensive blind spot, and a cheaper model is what makes closing it practical.
Safety, dual-use cyber and invite-only red-team access
SpaceXAI built Grok 4.7 with an entirely new safeguard stack and calls it the strongest model the company has tested on refusals and jailbreak resistance. It tops LatchBio's biosafety benchmark at 62.4%.
On HackerBench v0.3, SpaceXAI's own benchmark for risky and malicious cyber tasks, Grok 4.7 allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. SpaceXAI has started giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defence research.
Low refusal rates on legitimate security work is the number security-adjacent teams should note, and it lands in a year when agentic misuse has become a live reporting topic. For context on that direction of travel, see my read on agentic malware in Anthropic's threat intelligence report.
The counter-view on Grok 4.7
Three reasons not to re-platform this quarter.
- Single-provider risk: pinning a measurement programme to one vendor's model means your historical data set is only as stable as that vendor's release schedule and pricing. Model churn is fast, as I set out in what 261 releases say about model shelf life.
- Benchmark-to-workflow gap: CursorBench 4.0, DeepSWE v1.1 and Terminal-Bench 4.0 measure coding and terminal work. None of them measures whether a model summarises your category accurately or cites your brand. Run your own eval set.
- The speed claim needs testing: SpaceXAI headlines "twice as fast", and separately sells a fast variant at twice the output speed for twice the price. Benchmark the standard tier against your own latency needs before you budget for the premium one.
What to do with Grok 4.7 this week
- Pull your current cost per 1,000 measurement prompts and re-run the maths at $2 input and $6 output per million tokens. If the number halves, raise sampling frequency before you raise budget.
- Try Grok Build for free on one real job, ideally a repetitive content QA pass, and compare the output against your existing model.
- Add one conversational surface to your monitoring scope, even manually, so you have a baseline before the next model release moves the answers again.
- Hold your 2027 tooling budget loosely. Cheap frontier inference pressures OpenAI, Anthropic and Google on price, and I would not sign a two-year commitment at today's rate card.
Grok 4.7 is not the highest-scoring model on most of the tables SpaceXAI published, and the company does not claim it is. It is a frontier-class model at $2 and $6 per million tokens, announced on 21 September 2026 and available now. Treat the pricing as the news, and let it change how often you measure rather than what you measure.
This Grok 4.7 analysis was published 21 September 2026 and last updated 21 September 2026.
Tags