A week-old prediction from Elon Musk — Grok 4.7 should be roughly on par with Opus 5, not 5.1 — met reality on September 21 in a quiet, identical pricing, identical speed launch that Matthew Berman dissects across 2,819 words of charts and caveats. xAI's blog frames Grok 4.7 as its most capable model for coding and knowledge work : it stays on task longer, checks its own work more carefully and ships with the best-calibrated safeguard stack yet . Under the hood the change is concrete: a larger base model , a longer reinforcement-learning run weighted toward harder, multi-hour tasks , and native Grok Bot harness understanding , plus better document and slide generation. Pricing is deliberately frozen — $2 per million input tokens, $0.50 cached, $6 output , double for the fast variant — and the delay has a disarming explanation straight from Musk: the team may have penalized response length too much or something in RL , so the model still gives up too early on hard tasks and does not verify rigorously enough ; a few more days of cooking was the fix.
Cost and Intelligence Density: Reading CursorBench
CursorBench 4.0 is the long-horizon IDE-agent test that buyers now watch, and the disclosure that xAI now owns Cursor hangs over the chart. Berman plots score on the Y axis, average cost per task on the X axis ; the sweet spot is top-right — high performance, low bill. At low effort Grok 4.7 sits at 33%, at extra-high it climbs to 46.3% , one of the largest deltas outside GPT-5.6 Soul in the video. At the frontier it lands just behind Opus 5 on max thinking , yet at about half the cost to run the benchmark . A second view swaps cost for average output tokens per task on X; lower is better because it signals intelligence density — how much smarts fits per token when price per million is factored in. Through that lens Grok 4.7 nearly overlaps Opus 5 ; the lowest-effort point scores poorly but sips tokens, so it stays cheap. The true head-to-head emerges in the middle: Opus 5 in blue, Grok 4.7 in white, GPT-5.6 Soul edging in, with Fable 5.1 the outright winner on score at roughly the same token budget , but because its per-token price is several times higher , its cost per completed task ends up much higher . The takeaway is explicit: score alone does not decide — score divided by bill does .
The third cut of the same chart is steps per task , where GPT-5.6 Soul looks the most efficient — matching Berman's lived experience that Soul finds the most direct shot to a solution . Fable 5.1 remains near the top. Grok's family is consistent just behind the frontier on steps and tokens: cheap and quick at low effort, pricier but more accurate at extra-high. The model card and Cursor docs formalize the choice: 256k standard, 500k long context , with pricing that doubles for standard and triples for fast above 256k , triggered only by input length . That rule turns an architecture decision into an economic one : scaling context is not free, so strategy must weigh whether the window is truly needed.
The coding story then moves to DeepSWE v1.1 (DeepSuite) — historically the closest proxy for how engineers feel about models in production. The snapshot: Grok 4.7 71%, GPT-5.6 Soul 72.7%, Fable 5.1 70% ; on its own Grok looks ahead. Berman flags the curation : xAI's blog omits Astra from that table. When Astra re-adds itself, the ranking flips — Astra Max 74.1% takes the lead , making DeepSWE's true summit change hands. AA Briefcase (multi-hour office work) shows Grok 4.7 and Fable 5.1 Max essentially tied, while Terminal-Bench 4.0 separates cleanly: Grok 4.7 38% versus Fable 5.1 57.9% and Astra 58.2% — a gap that matters because terminal skill determines how quickly an agent moves . Legal work inverts the hierarchy: Grok 4.7 19.6% dominates the field , building on a family strength ( 4.6 at 15.8% ); Soul 2.5%, Fable 6.7%, Astra 5.4% fall away, yet Muse Spark 1.2 at 42% beats them all . HealthBench Professional is a wash across the board, Electrical Engineering Bench at 66% (card) puts Grok second only to Astra — remove Astra and Grok would win. Pattern: Grok is frontier-adjacent on office and legal, a tier behind on heavy terminal and coding .
Price, Context Window and the Economics of the Race
xAI builds pricing deliberately as one step off the frontier plus abundant supply . Grok 4.7 at $2/$6 sits against Opus 5 $5/$25 and Fable 5.1 $10/$50 — roughly 2× for Soul, 5× for Fable and Astra . The card and Cursor docs make the long-context multiplier explicit, yet the business logic is over-supply : xAI over-invested in GPU capacity and could not generate enough frontier demand to burn it, so lowering price creates usage . Berman links this to the Gavin Baker note: enterprise tokens are 62% open-weight versus 38% closed-weight , because open models are cheaper, more controllable, more private and customizable. Even so value still accrues mainly to OpenAI and Anthropic — since a token's flops, memory and watts are the same whether open or closed, the margin simply shifts to the infrastructure layer . The lesson is blunt: you cannot charge frontier prices without the frontier answer ; some industries will pay 5× for the best answer, others want good-enough automation, done cheaply . Whether a lab packs more intelligence into each token or rations output to spread capacity decides the margin story.
GDPval, Briefcase and a New Safety Stack
The reality check for knowledge work comes as two Elo tables: GDPval-AA Fable 5.1 1735, Grok 4.7 1695, Astra 1542 ; AA Briefcase Fable 1678, Grok 1657, Opus 5 behind . Why xAI shows Astra in GDPval but hides it in DeepSWE becomes telling — transparent where strong, quiet where weak . The blog's other headline is the entirely new safety stack : strongest refusal and jailbreak resistance tested to date , Q* prompter , and leadership on dual-use domains (cybersecurity, biology) for both benign utility and safe refusal . Numbers: LatchBio biosafety 62.4% on top , LatchBio Capabilities 44.5% at xhigh , HackerBench v0.3 lets only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. xAI notes invite-only access for select cyber partners to red-team capabilities. Mechanism: 11 capability benches (RNA expression analysis, epigenomic profiling, variant analysis, therapeutic discovery, preclinical pharmacology, pathogen surveillance, functional inference) averaged, with BioSecBench-Refusal/Surveillance/Function probing hidden biological hazards and raw pathogen data. Result: Grok 4.7 keeps general biology knowledge while measurably stepping up dual-use capability , but those measures are capability without safeguards — product filters sit on a separate layer.
The launch footnote tests the big claim. Elon's August 12 post that Grok 4.7 will exceed all current models now reads against reality — close to Fable 5 yet not Fable 5.1 , with Fable already present at the time — and Berman calls it very typical Elon: aggressive schedules, aggressive forecasts , with a long-run bet not to bet against him. Musk's delay note matches the tone: Grok 4.7 requires a few extra days to mature, the team may have over-penalized reply length during reinforcement or another factor played a role — the throwaway “ or something ” draws a wry smile in the video, yet the diagnosis is crisp: quits too early on hard tasks, insufficiently rigorous in self-checks . Availability is immediate: Cursor and Grok Build today, plus Grok API, third-party harnesses and model routers , identity grok-4.7 . Berman's personal through-line is about competition itself: more contenders is good for users ; Grok 4.7 may or may not hit the frontier at 4.8 or 4.9, but for now it is close enough and strikingly cost-effective .
External validation comes from Artificial Analysis Intelligence Index v4.3.2 : Fable 5.1 first, Astra second (tied), Opus 5 third, Muse Spark 1.3 Max 48 fourth, Grok 4.7 (xhigh) 46 fifth — GLM-4.7 and Kimi K3 close behind as open-weight peers. The index blends ten evaluations: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam and more. The +2-point lift over Grok 4.6 looks modest, but +111 Elo on AA-Briefcase to 1657 and +90 on GDPval-AA to 1695 mark long-horizon agent work. The caveat confirmed by the blog: gains come with higher token use — ~81k output tokens per task for Grok 4.7 versus 38k for 4.6, 60k for Muse Spark 1.3, 27k for Astra . So intelligence index alone does not tell the cost story; cost per task and speed must be read alongside . Berman noticing that Grok 4.7 is selected on the cost-per-intelligence view yet the dot does not appear, while 4.6 sits around $2,300 total to run the index , sets the expectation that 4.7 will live in the same low-cost band .
The context-window footnote also belongs in the decision table: Grok 4.6 and 4.7 offer 256k standard, 500k long , almost every other frontier offers 1M . For sub-250k jobs the difference may be invisible, but for very large repos, long legal dossiers or hundred-page reports the ceiling dictates. The billing nuance matters in practice: up to 256k at base rates, above that 2× (standard) or 3× (fast) — long-context pricing triggers on input length only . That ties cost arithmetic to context strategy : compress, chunk or cache when possible to keep the bill down. It is the same layer where architecture becomes economics .
Field notes are mixed yet instructive. Musk's ‘as of an hour ago xAI places third after Anthropic and OpenAI for agentic coding’ sets the competitive tone; Bobby's Grok 4.7 versus Kimi K3 demo — called ‘terrible as hell’ in the video — is noted by Berman with the proviso of settings and harness differences before generalizing. Automation bridges like Zapier (9,000+ apps, MCP server, references Cursor/Nvidia/Samsung/Dropbox/Shopify, Gmail-to-Asana flow) show where the model shines as a price-performance lever for knowledge-work automation , not as hype but as workload fit . The bottom line stays plain: Grok 4.7 is not the absolute frontier, one step behind; yet at half the price and same speed it turns ‘good enough’ into ‘highly sensible’ for teams that need automation . Berman closes on whether 4.8 reaches the frontier is unknown, but the competition itself already pays off — the natural next hope being a Grokbot update to 4.7 .
Cost Per Task Axis for Coding Frontier
- Grok 4.7$3 avg
- Opus 5$10 avg
- Fable 5.1$15 avg
- Muse 1.3 open~$1.5
| Topic | Status |
|---|---|
| Grok 4.7 price | $2 / $6 same speed, 500K long |
| Coding peak | DeepSWE 71%, Terminal 38%: behind |
| Office & legal | GDPval 1695, legal 19.6%: frontier-adjacent |
| Model | Input / Output | CursorBench |
|---|---|---|
| Grok 4.7 | $2 / $6 | 46.3% xhigh |
| Opus 5 | $5 / $25 | ~47–48% |
| Fable 5.1 | $10 / $50 | ~52%+ |
| Muse Spark 1.2 | open-weight | 42% legal |
Key moments
AI commentary
"My first instinct watching was not to be fooled by the top-right corner. Grok 4.7 nearly matching Opus 5 on CursorBench looks like a win at first glance; I chose to run every chart through the **cost, token-density and context-window** filter, because reading only the score without the price would mislead anyone picking a model for real work."
AI assessment
Steel-manning the bull case first: xAI freezing price while stretching RL for longer horizons genuinely shifts the cost-performance frontier , especially for office and legal production where long-context generation dominates; for a procurement desk that optimizes cost per completed task , paying 5× for a ~5-point score edge from Fable 5.1 may not be rational, and the safety stack — LatchBio 62.4% and HackerBench 3.3% — adds a corporate risk filter plus.
The limit shows in method and presentation. xAI owning Cursor creates a perception of selection bias; dropping Astra from the DeepSWE table strengthens the cherry-picked read; the Terminal-Bench gap from 38% to the 58% band is material for heavy agentic coding, and output tokens more than doubling from 38k to 81k inflates the bill per task; the 500k versus 1M context ceiling matters for very large dossiers. Some scores are harness-specific to Grok Build and await independent replication.
On verification the picture is mixed but testable: the x.ai blog and model card are consistent on price, context and RL narrative; Artificial Analysis Elo and token counts line up on independent measurement ( Briefcase 1657, GDPval 1695, ~81k tokens ); llm-stats cautions about cross-effort comparisons to keep the headline from inflating. Still, ‘best’ is benchmark- and harness-sensitive , and the Muse Spark 42% legal outlier shows a single-model niche win cannot be generalized.
The practical read is selective: for teams with high automation volume, medium error tolerance and contexts within 256–500k , Grok 4.7 is a striking lever and may be the price-performance champion for legal production and office documents . For terminal-heavy, many-step agentic coding and million-token depth , Fable/Astra/Opus remain a tier ahead ; those who need the best answer will still pay 5×. Decide on dollars and token density per task, not on the scoreboard alone .
Sources
8 links; 3 of them also cited by 6 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Matthew Berman: Did Elon catch up? (Grok 4.7 is here)
- @x.ai https://x.ai/news/grok-4-7
Also cited by: The AI Price War: Opus 5.5, GPT-6 Sol and the Quiet Bottleneck of the Agent Era · How OpenAI's Early Self-Improvement Push, DeepSeek's Sandbox Machine and Grok 4.7 Redraw the AI Race · Grok 4.7: Elon's Big Promise Tested on the Leaderboard and the Wallet · Grok 4.7 in Three Heavy Builds: Is One-Fifth the Price Enough Against Fable 5.1 and Astra?
- @media.x.ai https://media.x.ai/v1/website/4p7card-5eccc980.pdf
- @artificialanalysis.ai https://artificialanalysis.ai/articles/benchmarking-grok-4-7
Also cited by: The AI Week That Packed World Models, Open Weights and Data Centers Into Orbit
- @llm-stats.com https://llm-stats.com/blog/research/grok-4-7-launch
Also cited by: Grok 4.7 Lands: Same Price, Longer Horizons — Field-Tested Across 20 Builds
- @cursor.com https://prod.cursor.com/docs/models/grok-4-7
- @claude.com https://platform.claude.com/docs/en/about-claude/pricing
- @artificialanalysis.ai https://artificialanalysis.ai/models/releases/grok-4-7
grok 4.7 · xai · cursorbench · deepswe · terminal-bench · ai pricing · context window