Retour au fil

Grok 4.7 : la grande promesse d'Elon mise à l'épreuve des benchmarks et du prix

Grok 4.7 promettait de battre tous les modèles : le verbatim de SpaceXAI, les chiffres d'Artificial Analysis, le coût par tâche et l'expérience Grokbot passés au crible par Theo - t3.gg.

Importé dans Nodesdaily : (UTC+03:00)
Voir sur YouTube — jLgpzgpsWPc
Options de lecture

La lecture vocale n’est pas disponible dans ce navigateur.

Loupe à concepts

Choisissez un terme technique de cette vue pour lire sa définition générale, un exemple pédagogique et son usage dans l’article.

Aucun terme de notre glossaire n’a été trouvé dans cette vue. Le glossaire ne couvre pas encore tous les termes.

A month of hype built on a SpaceX data promise

For more than a month Elon framed Grok 4.7 as the big one, the successor that would beat every current model, with SpaceX's training corpus as the supposed edge. Even when 4.5 landed, the message was that 4.7 would be the leap where nothing beats it at real-world coding. Theo opens on that promise because the punchline is awkward: on several public boards 4.7 trails 4.6, which makes the earlier certainty feel like marketing rather than measured guidance. In my reading the issue is less that a model missed and more that the expectation was invited to be absolute; when the board says otherwise, trust erodes faster than any single score.

Why this video is about what numbers hide, not what they show

The thesis is that benchmarks no longer describe reality and the numbers we quote are the wrong proxies. Theo is explicit that he actually likes Grok 4.7, putting it among his favorites this year, yet finds the cloud of useless stats around the launch frustrating. I share that split view: a product can be enjoyable to use while the way it is measured is misleading. The video therefore oscillates between the official post and independent probes, trying to separate feel from figure. That structure matters because it keeps affection for the model from becoming cover for the messaging.

SpaceXAI's launch note calls Grok 4.7 its most capable system for coding and knowledge work, saying it stays longer on hard tasks, double-checks its own output more carefully, and ships with the best-calibrated safety stack so far. The note adds that price and speed stay identical to Grok 4.6, which is presented as proof of competitiveness. That framing collides with what Elon sketched months ago: 4.6 as a 1.5-trillion-parameter interim step in early August, 4.7 as a fresh 2.1-trillion base a few weeks later that would be stronger in every way, only marginally slower but more token-efficient. On paper those two stories cannot both be true, and the video shows which one the meter supports.

Artificial Analysis makes the conflict concrete. On its high-effort tier Grok 4.6 averages roughly 38 thousand tokens per task while Grok 4.7 jumps to about 81 thousand, above Fable 5.1. Soul sits near 29 thousand and Astra is in the same neighborhood, so Grok 4.7 is roughly doubling its predecessor. The note's claim of better token efficiency looks inverted when you look at cost per task on that same board. My take is simple: a larger base can explain higher capability, but it cannot explain higher capability at higher spend while claiming better efficiency; one of those adjectives has to be retired.

The request versus task sleight of hand

Cursor chief Michael Truell, now closely tied to xAI, pushed back by saying 4.7 uses about five percent more tokens for median requests and twenty to thirty percent more for the worst one percent, with a shift in what people ask. Theo argues that this reframes the unit of account in a way that flatters the model. The board uses fixed prompts, so distribution shift does not rescue the comparison, and in agentic work a single prompt can fan out into many requests. If the number of requests per prompt rises, cost per request looking modest does not keep cost per prompt modest, and Artificial Analysis is leaning more agentic over time, which amplifies the effect. Either the metric was chosen poorly or the rebuttal was. I found Theo's objection convincing because it ties the definition of price to the definition of work.

The dollar consequences are not abstract. Theo's own run of the board cost about twenty dollars for Grok 4.7 at list API rates, while no other model crossed ten dollars except one Astra run at 11.44, and every Opus 5 run stayed under a quarter of Grok 4.7's worst case. On Artificial Analysis the weighted cost stayed at 1.86 for Grok 4.6 and 3.26 for Astra while 4.7 hit 3.74 per task. The official sheet still lists two dollars per million in and six dollars per million out, but the fine print doubles price beyond two hundred thousand tokens and doubles again for the fast variant, which lands hard on normal use. It is the classic split between sticker price and bill; the sticker did not change, the bill did.

CursorBench is where the measurement story gets interesting. An earlier leak of benchmark data into training has been cleaned, so the current scores are more trustworthy. Astra is missing, which the video attributes to OpenAI blocking xAI from using its models in head-to-head posts and from sharing the harness with outsiders. Even with that hole, the picture is consistent: Grok 4.7 looks strong against Soul at most effort levels, but Fable low can still undercut Grok high on both score and price, which undermines the efficiency narrative. The only place Grok 4.7 is clearly efficient is its lowest-effort setting where it matches Soul high on tokens. The other notable detail is the Grokbot harness: the model is trained to understand that harness natively, which helps in Cursor and Grok Build even if it does not move the public board much. That native fit explains why users report smoothness that the board does not capture.

Outside the house charts, the air gets thinner

Switch to software and terminal work and the selection looks cherry-picked. On DeepSWE v1.1 Grok 4.7 beats Fable 5.1 but trails Soul; on Terminal-Bench the headline looks friendly until you see Fable 5.1 is twenty points ahead, close to a fifty percent lead. Legal and health evaluations like Harvey and HealthBench Professional put Grok 4.7 ahead of 4.6 and near other frontier names, but the set as a whole feels idiosyncratic. The video notes that GDPval still shows 4.6 ahead of Astra, a tell that older or narrower evaluations flatter the house. My reading is that the launch post needed to showcase a strange mix precisely because no compact public board captures what users like about the model; the curation is honest about that difficulty and less honest about the odds.

Independent aggregates are cooler. Artificial Analysis's new coding-agent index gives Grok 4.7 a 56 while Codeex and Claude Code with Fable and GPT-6 Astra sit near 62, and Muse Code and Muse Spark 1.3 are just below. On the older general index Grok 4.7 sits just above GLM-5.3 Max and below Muse Spark 1.3, well behind the two clear leaders Astra and Fable 5.1. Frontier Code from Cognition is called out as noisy, with non-monotonic curves by effort level; Grok 4.7 even slips below 4.6 in aggregate because it is said to overthink easier items while staying strong on hard ones. xAI's own note admits that trade. I think a six-point corridor looks small on a plot but feels large in a work session because it proxies how often you must intervene; as that rate falls, the usable horizon expands more than headline difficulty.

Pricing on subscription completes the squeeze. The sheet price of two and six dollars per million, with a fast option at double, is only half the story once the tariff doubles beyond two hundred thousand tokens and the fast tier adds another multiple for everyday use. Co-host Ben Davis, a noted Grok defender, flags that forty dollars of metered use burned eight percent of a weekly allowance on a three hundred dollar plan, implying a tight weekly budget when rivals offer eight to twelve thousand on two hundred dollar plans. Previously the three hundred dollar tier felt effectively unlimited for Grok; now it feels rationed without the premium feel of Fable or Astra. When a model sits in the Soul/Opus performance corridor but carries frontier-tier throttling, the economics invite comparison shopping.

Frontend ability is where the video is most blunt. Given the same prompt and environment that produced a decent-looking game from Astra, Grok 4.7 delivered one of the poorest frontends seen in well over a year, with reversed fish and submarine motion and broken cursor tracking that even the author calls nauseating. Anthropics design leadership is reaffirmed and Astra is at least steerable, but Grok 4.7 is judged harshly on presentation ELO, which Artificial Analysis also marks as regressing versus 4.6. Benchmarks rarely reward taste and coherence, yet those qualities decide whether a generated app feels shippable; here the board and the eye agree negatively.

A late note explains why the model slipped a week. The team believes it over-penalized length in reinforcement learning, so the model quit early on hard tasks and did not check itself rigorously enough. The final week was spent pushing it to tolerate long outputs and to verify more, which likely erased the earlier efficiency boast. Field reports of strange loops fit that picture: a Grok Build session spiraling into an infinite build monologue about portals and infinite multiverses, and a K-pop translation path on X surfacing what looks like internal instruction text where a simple translation was expected. Both feel like side effects of the same late tuning that made the model more persistent but also more verbose and more leaky in edge integrations.

What keeps the author engaged is the floor rather than the ceiling. Grok 4.7 does not beat Fable or Astra at their peaks, but it stays on track longer, digs deeper, and sometimes surfaces issues that the other two miss. To probe that, Theo ran a small study on the T3 codebase asking several models to propose improvements, with Fable 5.1 and Astra acting as judges. Astra found the most novel and best-vetted suggestions, Grok 4.7 came just behind on viability, Fable 5.1 offered fewer but high-quality ideas and vetted less. Cost again tells the warning: cheapest Grok 4.7 runs match Astra/Fable, worst runs double them because the model can keep verifying beyond diminishing returns. That pattern mirrors Grok 4.6's habit of buying capability with tokens at the expense of latency and spend.

The closing demonstrations lock the verdict. The Fish Slop 3D scene, rendered from one shared prompt, looks solid from Astra and broken from Grok 4.7, with direction, speed and cursor logic wrong, and it took roughly an hour and a half to finish, which delayed filming. Combined with the earlier token and subscription notes, the author's earlier affection for Grok 4.5's speed and cheapness now reads as nostalgia; 4.6 and 4.7 have traded that charm for persistence without frontier-level payoff. The conclusion is that Grok 4.7 is a stepping stone awaiting more reinforcement learning for efficiency, hard to recommend versus Codeex or Claude unless you are already invested in Cursor's cloud. Elon was right that it would be interesting, just not in the way the hype promised, and the next iteration will need to reconcile performance with price to matter.

Visualization: nodesdaily AI

Commentaire de l’IA

"Je partage la double lecture de Theo : Grok 4.7 est agréable à utiliser, mais le récit autour des chiffres est trompeur; aimer le modèle n'oblige pas à croire le graphique."

Évaluation de l’IA

To steelman xAI, the strongest version of its defense is that in live traffic the median request is only about five percent pricier and even the tail is contained at twenty to thirty percent, and that with Grokbot the system feels smoother and needs less hand-holding even if it fans out into multiple requests. I believe that picture is sincere, because production traffic is not benchmark traffic and agentic use can inflate cost per prompt while improving experience. The problem is not that the defense is false but that it answers a per-task benchmark with a per-request metric; the two denominators do not match, so the reassurance misfires against the chart that actually flagged the cost.

The gaps remain clear. First, method: fixed-prompt boards measure cost growth directly, and a shift in what users attempt does not rescue a board built on static prompts. Second, selection: the launch set spotlights the strongest arenas while leaving dimmer corners dark, like a twenty-point deficit to Fable on Terminal-Bench and trailing Spark on the general index. Third, price transparency: sticker price is flat yet the bill is not, with a doubling beyond two hundred thousand tokens and another multiple for the fast variant, plus weekly quota burn of eight percent for forty dollars of use. Fourth, side effects: late reinforcement that pushed for longer, more verifying behavior also produced loop and translation leakage anecdotes, raising predictability as well as efficiency questions.

On provenance and checkability the record is mixed. Some scores have survived outside audit, including the cleaned CursorBench and Artificial Analysis's re-weighted agentic index, but the structural absence of Astra from house charts leaves a hole, and the invisibility of Fable 5.1 on DeepSWE at the official site invites caution. Safety claims score well on LatchBio and HackerBench, yet leaked instruction text in a consumer translation path suggests integration discipline is not uniform. I would read every score alongside its token and time cost versus the model's own predecessor, not just versus rivals, and treat leadership claims as cost-conditioned rather than absolute.

My practical take is this: Grok 4.7 is worth trying if you value an agent that stays on track longer and digs deeper, especially inside Cursor or Grok Build where harness fit matters. If your priority is speed and low invoice variance, Fable 5.1 and Astra still offer a tighter value shape, and if you need polished frontend output, Grok 4.7 currently lags. If you are quota-sensitive, budget for the beyond-two-hundred-thousand multiplier and the fast variant multiple, otherwise you risk parity in the best case and roughly double in the worst case per task. For me this release is not a frontier break but a stepping stone; without a follow-up tuning that restores some efficiency, the price-for-performance case struggles to justify a switch.

Sources

6 liens ; 2 d’entre eux sont aussi cités par 6 autres articles. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

grok 4.7 · xai · benchmarks

Suivre le sujet

Avant cet article

Un court ordre de lecture des articles antérieurs reliés à cet événement par un éditeur.

Preuves et sources

Consultez les passages autorisés, leurs versions et leur origine.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

Grok 4.7 : la grande promesse d'Elon mise à…