Back to feed

Grok 4.7 Lands: Same Price, Longer Horizons — Field-Tested Across 20 Builds

Released Sep 21 2026, xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing under 200k tokens while lifting long-horizon coding and agent scores; a field test across 20 builds shows a cheap, fast model that still needs caution on detail.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 6Rfk1P98q74
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

On Sep 21 2026 xAI shipped Grok 4.7 as a same-price upgrade; the video calls it Grock from SpaceX AI, but the canonical name is Grok from xAI , and host Julian Goldie tests it after 20+ builds on his Agent OS bench called Goldiebench . My read: this is not a revolution but a tougher training diet on the same commercial shape — same bill, longer persistence. Like reaching one floor higher on the same scaffolding, cost stays flat while endurance grows.

Same Price, Better Scores

The LLM Stats synthesis of launch notes shows Grok 4.7 keeping the same 500k context class: under 200k prompt tokens $2 input / $0.50 cached / $6 output per million, and $4 / $1 / $12 once you cross 200k for the whole request. Gains cluster on long-horizon benches: CursorBench 4.0 40.4 → 46.3% , DeepSWE 65.2 → 71.0% , Terminal-Bench 4.0 20.3 → 38.0% , EEBench 60 → 66.0% — all vendor-reported and awaiting independent replication. Like the same scaffolding reaching higher, a 10-page repo fits with less drift on the same invoice. The model still takes text + image in and text out; pretraining cutoff is June 2026 with supplements through August 2026 .

The lift comes from a harder curriculum, not a new price tag. The practical lever is reasoning effort : four levels — low, medium, high, xhigh — default high . Like choosing sandpaper grit, low is fast and cheap, xhigh is slow and deep; most headline scores are reported at xhigh , so re-measure token-per-task at each level. A simple summary stays cheap on low, while a multi-step agent loop burns more tokens on xhigh; the default budget does not carry over.

Context stays at 500k , same league as Grok 4.6. The launch post mentions a fast variant at 2× price and 2× output speed, yet the docs at publish time list only the grok-4.7 id, no separate fast id. In practice the 200k cliff is a switch: short prompts stay on the standard rate, long prompts flip the entire request to the higher tier. Start at high , move to xhigh only if the bottleneck is duration, and meter tokens per level — default budgets do not carry over. The 500k window looks big, but price also doubles beyond 200k; plan both dimensions together for long chats.

Field Test on Goldiebench: From Twilight Veil to Neon City

On Goldie's bench the harness is called Goldiebench ; you can flick between Grok 4.7, 4.6 and 4.5 and preview every build in the shared workspace . The showcase Twilight Veil is a 3D world game: smooth play, easy setup, decent atmosphere, but — as the host says — not near Astra 6 class. Neon City and a flight simulator follow the same pattern — they run, they are quick, yet lack the texture and vibe you get from rival flagships the captions call Fable 5.1. Like a quick sketch: proportions right, material and shadow still raw. Community reaction splits the same way: cheap and fast is loved, those chasing final varnish stay on hold.

The recurring observation is lack of detail . In coding tests the model ships the feature and passes checks, but the final varnish is thin; the feel rivals provide is muted here. The trade is speed and price: same prompts return in seconds, and input and output tokens read cheaper than flagship peers — the video spotlights tight scores like 46.3 vs 41.7 for that reason (names in captions are ASR-noisy; the verified table is Grok 4.7 vs 4.6). For builders, CLI via X and direct Agent OS wiring help: add the CLI, keep every artifact in the workspace. Ask for a clock widget then add a filter — the scaffold holds, ornament lags. This shows the harness's persistence matters more than raw model cleverness; the bench that clicks and remembers outruns the model alone.

Price, Access and Roadmap

Price is simple: same label, more endurance. $2 / $0.50 / $6 is generous for hobby loops; $4 / $1 / $12 beyond 200k means shortening the prompt saves money. Access is open day-one on Grok Build , Cursor , Grok API and router clouds; CLI via X is part of the pitch. The fast option lives in prose, not in a separate id, so plan latency lanes from your account's actual model list, not from marketing copy — copy and API can diverge. Like a flight, advertised speed and gate speed do not always match; measure in your own account.

The roadmap moves fast: Goldie expects Grok 4.8 soon with releases every few weeks; Gate reports xAI aiming to close pretraining of a 2.5T-parameter Grok 4.8 by week's end. One-sentence promo distill: the host's AI Profit Boardroom / Agent OS is a 3,400-member community with a 30-day roadmap , four coaching calls a week and a mission board that sells the bench more than the model. My take: Grok 4.7 is a fun, economical alternative for daily tinkering and mid-scale agent loops, especially if you want to skip a Claude bill — not perfect alone, human review stays mandatory for fine detail and high-stakes calls. The 20-build field test reminds you to try cheap and fast while remembering who to trust for the real job.

Visualization: nodesdaily AI

Grok 4.7 Headline Scores

  • CursorBench 4.046.3%
  • DeepSWE71.0%
  • Terminal-Bench38.0%
  • EEBench66.0%
Higher is better; all values vendor-reported.
DimensionSummary
PriceSame tag: $2 / $0.50 / $6 (<200k) → $4 / $1 / $12 (≥200k)
ScoresCursor 46.3%, DeepSWE 71%, Terminal 38%, EE 66% (vendor)
Field20 builds: Twilight Veil smooth, detail needs polish
BenchmarkGrok 4.7Grok 4.6Delta
CursorBench 4.046.3%40.4%+5.9
DeepSWE71.0%65.2%+5.8
Terminal-Bench 4.038.0%20.3%+17.7
EEBench66.0%60.0%+6.0

Key moments

  1. Intro — Grok 4.7 in the wild
  2. Benchmarks vs rival flagships
  3. Bench — switching 4.7 / 4.6 / 4.5
  4. Demo — Twilight Veil 3D world
  5. Observation — fast, detail light
  6. Access — CLI via X and Agent OS
  7. Roadmap — Grok 4.8 near

AI commentary

"For me Grok 4.7 is not a 'cheaper' story but a 'more persistent on the same bill' story; the 20 builds in the video prove it — fast, economical and fun to try, yet I still leave fine detail work to others."

AI assessment

In my view the strongest claim is also the one needing the most careful reading: more persistent on the same bill . To steelman it: if your bottleneck is truly long-horizon coding and agent loops, higher scores in the same 500k window are concrete value. Keeping the under-200k price flat makes trying cheap, and seeing fluidity across 20 builds on Goldiebench supports the thesis. A harness that clicks and remembers lets the same model finish more work — a generalizable win for small teams on legacy tools.

What is missing is methodology and risk. All numbers are vendor-reported ; as LLM Stats warns, Terminal-Bench and EEBench run on Grok Build , coupling score to harness. EEBench reads 66.0 on the card but 64.0 in the news table; CursorBench 4.0 is not comparable to the prior version. The fast variant lives in prose but has no separate id in docs. Names of rival models in captions are ASR-noisy (the video even writes SpaceX AI and Grock), so read the verified table, not the caption sentence. The cheap-token pitch can also hide rising token-per-task at high/xhigh .

I checked incentives and verifiability. While xAI bundles Grok 4.7 on SuperGrok and Cursor, outside coverage in the same window flags hallucinations and biased outputs ; a RAND commentary notes regulatory heat. Price and plan bundling shifts monthly; docs.x.ai and x.ai/pricing need a live check at decision time. The Grok 4.8 2.5T claim and week's-end pretraining close is single-source. Token-per-task and fast-variant details should never be taken from one source.

My practical take: for daily tinkering, rapid prototypes and mid-scale agent loops, Grok 4.7 is currently among the lowest-friction paths — fast, economical, fun to try , a solid alternative if you want to skip a Claude bill. For handing production data and high-stakes instructions to the cloud, do not skip human review , re-measurement on your own evals and prompt shortening around the 200k cliff. Yes for builders, not yet for everyone seeking a general assistant; trust the model id in your account, not the marketing line.

Sources

7 links; 2 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

grok 4.7 · xai · cursor · agent os · goldiebench · pricing

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…