Grok 4.7 is not an uncontested new champion. It is something more useful: a credible near-frontier model with strong coding and professional-work results at a relatively low API token price.
The Short Verdict
SpaceXAI calls Grok 4.7 its most capable model for coding and knowledge work. The company says it uses a larger base model, a longer reinforcement-learning run, harder long-duration tasks, better self-checking, and more training around the Grok Bot harness.
Those changes appear to matter most when the model is allowed to work for longer inside an agent harness. The independent Artificial Analysis evaluation gives Grok 4.7 an Intelligence Index score of 46, up two points from Grok 4.6 high, and a Coding Agent Index score of 56 with Grok Build. That coding result trails Fable 5.1, GPT-6 Astra, and Opus 5 in their native harnesses, but it is a substantial improvement over Grok 4.6.
The catch is efficiency at high reasoning effort. Artificial Analysis measured about 81,000 output tokens per Intelligence Index task, compared with 36,000 for Grok 4.6 high and 27,000 for GPT-6 Astra max. Grok 4.7's token rates are inexpensive; the model may still spend more tokens getting to the answer.
Watch Matthew Berman's Full Review
Source and credit: Did Elon catch up? (Grok 4.7 is here), published by Matthew Berman on 21 September 2026. His video treats the release charts as selected evidence, checks them against Artificial Analysis, and concludes that Grok 4.7 is close to the frontier rather than definitively on top.
What Changed in Grok 4.7
The official release describes four connected changes: a larger base model, a longer reinforcement-learning run, more difficult long-horizon tasks, and better understanding of how to operate inside the Grok Bot coding harness. It also says the model works longer, checks its own work more carefully, and uses a new safeguard stack.
This is important because a model and a harness are a system. Planning, file access, terminal tools, subagents, retries, and verification can change the result as much as the raw model. A score from Grok Build is evidence for Grok 4.7 inside Grok Build; it is not automatically the score every third-party agent will reproduce.
SpaceXAI lists access through Grok Build, Cursor, its API, coding harnesses, routers, and cloud platforms. Artificial Analysis reports configurable reasoning effort from low through xhigh and a 500,000-token context window.
Read the Official Benchmark Table as Selected Evidence
The following numbers come from SpaceXAI's release page. They are useful, but they were selected and reported by the model provider. Effort settings and harnesses differ, and the set does not establish a universal ranking.
| Official evaluation | Grok 4.7 | GPT-5.6 Sol | Fable 5.1 | What it suggests |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3 | 41.7 | 51.8 | Competitive in Cursor, still behind Fable. |
| DeepSWE v1.1 | 71.0 | 72.7 | 70.0 | Close three-way coding result. |
| AA-Briefcase | 1657 | 1487 | 1678 | Near-frontier professional work. |
| Terminal-Bench 4.0 | 38.0 | 37.3 | 57.9 | Good terminal result; Fable leads clearly. |
| Harvey Legal | 19.6 | 2.5 | 6.7 | A notable provider-reported specialty win. |
| HealthBench Pro | 56.7 | 60.5 | 62.1 | Competitive, not leading. |
Matthew Berman explicitly warns that the launch presentation is cherry-picked. He also adds comparison context that is not part of SpaceXAI's original chart. Keep those layers separate: provider data, creator analysis, and independent evaluation answer different questions.
The Independent Result Is Stronger Than the Hype Needs
Artificial Analysis evaluated Grok 4.7 at xhigh reasoning effort. Its Intelligence Index rose to 46, with the biggest gains concentrated in agentic knowledge work. AA-Briefcase reached 1657 Elo, 111 points above Grok 4.6 high, while GDPval-AA reached 1695, up 90 points.
With Grok Build, the Coding Agent Index rose from 47 to 56. The component results were 73% on DeepSWE v1.1, 33% on Terminal-Bench 4.0, and 63% on SWE-Atlas-QnA. Artificial Analysis notes that these are native-harness results. Its separate Intelligence Index uses a standardized harness, so the two scores should not be blended into one ranking.
Not every capability improved. Artificial Analysis reports smaller gains outside agentic knowledge work, including a 4.5-point Terminal-Bench improvement and a three-point GDP.pdf gain, alongside regressions on AA-LCR and AutomationBench-AA. Its hallucination test improved from 34% to 29%, while measured accuracy remained broadly flat at 47% versus 48% for Grok 4.6 high.
That pattern supports a narrow conclusion: Grok 4.7 is meaningfully better at sustained analytical and agentic work. It is not evidence that every short answer, automation, design, or coding task improved.
Low Token Price Is Not the Same as Low Task Cost
| Economic input | Reported value | Decision implication |
|---|---|---|
| Standard input | $2 per 1M tokens | Low entry price for a frontier-class proprietary model. |
| Standard output | $6 per 1M tokens | Attractive rate, but long reasoning can still accumulate cost. |
| Cached input | $0.50 per 1M tokens | Useful for repeated instructions and stable project context. |
| Fast variant | About 2x output speed at 2x price | Buy latency only when response time affects the workflow. |
| AA output use | About 81k tokens per task at xhigh | Compare completed-task cost, not token rates alone. |
| AA task estimate | $3.74 per Intelligence Index task | A benchmark estimate, not a quote for your workload. |
A model can be cheap per token and expensive per accepted result if it reasons for longer, retries tools, produces oversized artifacts, or needs human correction. Conversely, a higher-priced model can be cheaper when it reaches the acceptance bar in one pass. Track successful tasks, review time, and reruns beside the API bill.
Speed also depends on prompt length, reasoning effort, provider, and measurement method. Artificial Analysis reports roughly 188 output tokens per second for long-prompt measurements and about 7.1 minutes per Intelligence Index task. Its model comparison pages use other summary measurements. Treat speed as a workload test, not one permanent label.
Where Grok 4.7 Belongs in a Model Router
- Test first: multi-file coding, terminal work, research-heavy documents, spreadsheets, slide deliverables, and legal-work triage with expert review.
- Route selectively: high-volume jobs where $2/$6 token pricing matters and xhigh reasoning is not required on every request.
- Keep alternatives: Fable 5.1 for stronger native-agent and terminal results, GPT-6 Astra for workflows where its lower measured output-token use or computer control matters, and Opus 5 for tasks where it wins your own quality gate.
- Check context needs: 500,000 tokens is large, but it may not fit unusually long repositories or archives without retrieval, summaries, or partitioning.
- Do not infer safety: SpaceXAI reports stronger refusal, jailbreak, and dual-use safeguards. Those are provider-reported results, not an independent assurance for a specific deployment.
Run a Five-Task Evaluation Before Switching
- Choose five real tasks: one bug fix, one feature, one terminal investigation, one long research deliverable, and one task where your current model often fails.
- Freeze the environment: use the same repository state, tools, permissions, and acceptance tests.
- Test two effort levels: compare high and xhigh instead of assuming maximum reasoning is always economical.
- Record the whole task: elapsed time, input and output tokens, tool calls, retries, human review minutes, and whether the first result passed.
- Adopt by route: assign Grok 4.7 only the task classes where it improves accepted cost, speed, or quality.
The model earns a place when it reduces cost per approved outcome. A benchmark lead is interesting; a repeatable workflow win is the buying decision.
Video Chapters
| Time | Chapter | Time | Chapter |
|---|---|---|---|
| 00:00 | Introduction | 01:11 | Performance benchmarks |
| 03:35 | Zapier sponsor | 05:06 | Deeper model testing |
| 08:44 | Pricing and positioning | 10:38 | Safeguards and usage |
| 11:53 | Development context | 13:09 | Artificial Analysis |
| 15:06 | Criticism and limitations | 16:33 | Final thoughts |
Sources and Links
- Matthew Berman: Did Elon catch up? (Grok 4.7 is here)
- SpaceXAI: Grok 4.7 official release (21 September 2026)
- Artificial Analysis: Benchmarking Grok 4.7 (21 September 2026)
- Artificial Analysis: Grok 4.7 model page (live measurements may change)