AI Model Reviews

Grok 4.7 Review: Near-Frontier Coding at a Lower Token Price

Grok 4.7 is not an uncontested new champion. It is something more useful: a credible near-frontier model with strong coding and professional-work results at a relatively low API token price.

The practical answer to “Did Elon catch up?” On selected coding, terminal, legal, and long-horizon knowledge tasks, SpaceXAI is now close enough to deserve a production evaluation. Across the whole frontier, Fable 5.1, GPT-6 Astra, and Opus 5 still lead important evaluations. Grok 4.7's best argument is price-performance, not universal superiority.

The Short Verdict

SpaceXAI calls Grok 4.7 its most capable model for coding and knowledge work. The company says it uses a larger base model, a longer reinforcement-learning run, harder long-duration tasks, better self-checking, and more training around the Grok Bot harness.

Those changes appear to matter most when the model is allowed to work for longer inside an agent harness. The independent Artificial Analysis evaluation gives Grok 4.7 an Intelligence Index score of 46, up two points from Grok 4.6 high, and a Coding Agent Index score of 56 with Grok Build. That coding result trails Fable 5.1, GPT-6 Astra, and Opus 5 in their native harnesses, but it is a substantial improvement over Grok 4.6.

The catch is efficiency at high reasoning effort. Artificial Analysis measured about 81,000 output tokens per Intelligence Index task, compared with 36,000 for Grok 4.6 high and 27,000 for GPT-6 Astra max. Grok 4.7's token rates are inexpensive; the model may still spend more tokens getting to the answer.

Watch Matthew Berman's Full Review

Source and credit: Did Elon catch up? (Grok 4.7 is here), published by Matthew Berman on 21 September 2026. His video treats the release charts as selected evidence, checks them against Artificial Analysis, and concludes that Grok 4.7 is close to the frontier rather than definitively on top.

What Changed in Grok 4.7

The official release describes four connected changes: a larger base model, a longer reinforcement-learning run, more difficult long-horizon tasks, and better understanding of how to operate inside the Grok Bot coding harness. It also says the model works longer, checks its own work more carefully, and uses a new safeguard stack.

This is important because a model and a harness are a system. Planning, file access, terminal tools, subagents, retries, and verification can change the result as much as the raw model. A score from Grok Build is evidence for Grok 4.7 inside Grok Build; it is not automatically the score every third-party agent will reproduce.

SpaceXAI lists access through Grok Build, Cursor, its API, coding harnesses, routers, and cloud platforms. Artificial Analysis reports configurable reasoning effort from low through xhigh and a 500,000-token context window.

Read the Official Benchmark Table as Selected Evidence

The following numbers come from SpaceXAI's release page. They are useful, but they were selected and reported by the model provider. Effort settings and harnesses differ, and the set does not establish a universal ranking.

Official evaluationGrok 4.7GPT-5.6 SolFable 5.1What it suggests
CursorBench 4.046.341.751.8Competitive in Cursor, still behind Fable.
DeepSWE v1.171.072.770.0Close three-way coding result.
AA-Briefcase165714871678Near-frontier professional work.
Terminal-Bench 4.038.037.357.9Good terminal result; Fable leads clearly.
Harvey Legal19.62.56.7A notable provider-reported specialty win.
HealthBench Pro56.760.562.1Competitive, not leading.

Matthew Berman explicitly warns that the launch presentation is cherry-picked. He also adds comparison context that is not part of SpaceXAI's original chart. Keep those layers separate: provider data, creator analysis, and independent evaluation answer different questions.

The Independent Result Is Stronger Than the Hype Needs

Artificial Analysis evaluated Grok 4.7 at xhigh reasoning effort. Its Intelligence Index rose to 46, with the biggest gains concentrated in agentic knowledge work. AA-Briefcase reached 1657 Elo, 111 points above Grok 4.6 high, while GDPval-AA reached 1695, up 90 points.

With Grok Build, the Coding Agent Index rose from 47 to 56. The component results were 73% on DeepSWE v1.1, 33% on Terminal-Bench 4.0, and 63% on SWE-Atlas-QnA. Artificial Analysis notes that these are native-harness results. Its separate Intelligence Index uses a standardized harness, so the two scores should not be blended into one ranking.

Not every capability improved. Artificial Analysis reports smaller gains outside agentic knowledge work, including a 4.5-point Terminal-Bench improvement and a three-point GDP.pdf gain, alongside regressions on AA-LCR and AutomationBench-AA. Its hallucination test improved from 34% to 29%, while measured accuracy remained broadly flat at 47% versus 48% for Grok 4.6 high.

That pattern supports a narrow conclusion: Grok 4.7 is meaningfully better at sustained analytical and agentic work. It is not evidence that every short answer, automation, design, or coding task improved.

Low Token Price Is Not the Same as Low Task Cost

Economic inputReported valueDecision implication
Standard input$2 per 1M tokensLow entry price for a frontier-class proprietary model.
Standard output$6 per 1M tokensAttractive rate, but long reasoning can still accumulate cost.
Cached input$0.50 per 1M tokensUseful for repeated instructions and stable project context.
Fast variantAbout 2x output speed at 2x priceBuy latency only when response time affects the workflow.
AA output useAbout 81k tokens per task at xhighCompare completed-task cost, not token rates alone.
AA task estimate$3.74 per Intelligence Index taskA benchmark estimate, not a quote for your workload.

A model can be cheap per token and expensive per accepted result if it reasons for longer, retries tools, produces oversized artifacts, or needs human correction. Conversely, a higher-priced model can be cheaper when it reaches the acceptance bar in one pass. Track successful tasks, review time, and reruns beside the API bill.

Speed also depends on prompt length, reasoning effort, provider, and measurement method. Artificial Analysis reports roughly 188 output tokens per second for long-prompt measurements and about 7.1 minutes per Intelligence Index task. Its model comparison pages use other summary measurements. Treat speed as a workload test, not one permanent label.

Where Grok 4.7 Belongs in a Model Router

  • Test first: multi-file coding, terminal work, research-heavy documents, spreadsheets, slide deliverables, and legal-work triage with expert review.
  • Route selectively: high-volume jobs where $2/$6 token pricing matters and xhigh reasoning is not required on every request.
  • Keep alternatives: Fable 5.1 for stronger native-agent and terminal results, GPT-6 Astra for workflows where its lower measured output-token use or computer control matters, and Opus 5 for tasks where it wins your own quality gate.
  • Check context needs: 500,000 tokens is large, but it may not fit unusually long repositories or archives without retrieval, summaries, or partitioning.
  • Do not infer safety: SpaceXAI reports stronger refusal, jailbreak, and dual-use safeguards. Those are provider-reported results, not an independent assurance for a specific deployment.

Run a Five-Task Evaluation Before Switching

  1. Choose five real tasks: one bug fix, one feature, one terminal investigation, one long research deliverable, and one task where your current model often fails.
  2. Freeze the environment: use the same repository state, tools, permissions, and acceptance tests.
  3. Test two effort levels: compare high and xhigh instead of assuming maximum reasoning is always economical.
  4. Record the whole task: elapsed time, input and output tokens, tool calls, retries, human review minutes, and whether the first result passed.
  5. Adopt by route: assign Grok 4.7 only the task classes where it improves accepted cost, speed, or quality.

The model earns a place when it reduces cost per approved outcome. A benchmark lead is interesting; a repeatable workflow win is the buying decision.

Video Chapters

TimeChapterTimeChapter
00:00Introduction01:11Performance benchmarks
03:35Zapier sponsor05:06Deeper model testing
08:44Pricing and positioning10:38Safeguards and usage
11:53Development context13:09Artificial Analysis
15:06Criticism and limitations16:33Final thoughts

Sources and Links

Common questions

Did Grok 4.7 catch up with the best frontier models?
It is close on several coding and agentic knowledge-work evaluations, but not uniformly ahead. xAI reports wins on selected tasks, while Artificial Analysis places it near the frontier and behind the strongest native coding-agent systems overall.
How much does Grok 4.7 cost?
SpaceXAI lists standard API pricing at $2 per million input tokens and $6 per million output tokens, with cached input at $0.50 per million. A faster variant offers roughly twice the output speed at twice the token price.
Is Grok 4.7 cheaper in practice?
Its token rates are low for a frontier-class model, but total task cost also depends on reasoning effort, retries, tool calls, and output length. Artificial Analysis measured about 81,000 output tokens per Intelligence Index task at xhigh effort, so low token rates do not guarantee the lowest completed-task cost.
What is Grok 4.7 best suited for?
The strongest evidence is for long agentic knowledge work, coding in Grok Build, terminal tasks, document-heavy office work, and legal-work benchmarks. Teams should still test their own repositories, tool stack, and acceptance criteria.
What is the Grok 4.7 context window?
Artificial Analysis reports a 500,000-token context window, unchanged from Grok 4.6. That is substantial, although some competing frontier models offer larger advertised windows.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call