Back to feed

Grok 4.7 in Three Heavy Builds: Is One-Fifth the Price Enough Against Fable 5.1 and Astra?

Pat Simmons' **4,474-word** one-shot test puts xAI's claim for **Grok 4.7** — a real rival to Fable 5.1 and Astra at a far lower price — to work across three heavy builds: an **Awwwards pixel clone, a scroll-driven 3D keyboard page and a Rainbow Road Mario Kart**. With **46% vs 51.6% on CursorBench**, **71% vs 70% / 72.7% on DeepSuite**, **$2 / $6 vs $10 / $50** per million tokens and **$7 vs $123 / $97** per award build, the sheet shows why cheap does not carry heavy code.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — x48xbDO6fKo
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

xAI opens with the line that Grok 4.7 is a real Fable- and Astra-class model for much less money , and Pat Simmons decides to verify it in a single frame with three heavy builds. The setup is blunt: the same three prompts are fired at the same time on three models — clone an Awwwards winner pixel by pixel, build a scroll-driven 3D page for a mechanical keyboard, and clone Mario Kart on Rainbow Road — and the scoreboard is triple: how fast it shipped, what it cost, and how many tokens it burned . The promise of the video is to show whether a five-fold price gap actually closes real build quality or just makes the table look cheaper. The rule is one shot from the start; no model gets a revision.

From Press Note to Boards: What CursorBench and the Bill Tell Us

The first act is a recap of the sparse press note. On the CursorBench 4.0 chart the Y axis is success, the X axis average cost per task; the sweet spot is top-right. Fable 5.1 at extra-high hits 51.6%, Grok 4.7 at its highest effort ~46% — inside four points — but the pricing lane flips: Grok 4.7 at extra-high $2 input / $6 output per million, Fable and Astra $10 / $50 . Simmons' emphasis here is that even when scores look close, the cost per completed task tilts clearly toward Grok, which is why the ‘cheapest by far’ label sticks to it. An honest footnote stays on the chart: xAI now owns Cursor , so the table carries a bias perception note; and GBT6 Astra is missing from the CursorBench 4.0 table , so the comparison is presented selectively.

The second layer fills in headline scores. On DeepSuite Grok 4.7 71%, Fable 5.1 70%, GPT-5.6 Soul 72.7% — Grok looks a touch ahead on engineering feel, at least in the slice the note chooses. Then multi-hour office work (AA Briefcase-like) scores run neck-and-neck across Fable, Soul and Grok; Harvey's legal-agent benchmark shows Grok strong, and clinical reasoning also puts Grok 4.7 above its 4.6 predecessor and near the Soul/5.1 line . Simmons reads the note as deliberately bright on knowledge-work outside code while letting the other side of the story hide: terminal-heavy tables like Terminal-Bench invert the picture , which his wrap-up slide makes explicit — many rows where Grok beats Fable/Astra on office and legal sit alongside the note that Astra stays a few points ahead only on Automation Bench .

Method: Three Prompts, One Shot and the Clone-App Pro Rig

Simmons' method is intentionally spare: a single bash command launches the same three builds on three models at once . Each prompt text will be shared in the post. The shared track is one-shot — no second try — which excludes the reality that iteration would close gaps in daily use , but it captures the raw frontier photo . The toolchain is standard: for the Awwwards clone the clone-app-pro skill is loaded — walk the code, take screenshots, run QA loops — so taste, navigation, vision and instruction-following are tested together. The ending is a blind reveal : columns A/B/C are opened one by one and matched to their owners, with cost and time written beside them.

The prompt for the first build is short but demanding: go to Awwwards, pick an award-winning site you genuinely find impressive and rebuild it pixel by pixel; load the clone-app-pro skill, QA your own work with screenshots. Simmons says this directly probes taste — what the model finds impressive, whether the pick is generic or distinctive — followed by browsing the site, perceiving visuals, loading the skill correctly and building the web . Requirements are minimal; the craft hides in how faithfully the model reconstructs an existing award language rather than inventing from emptiness. Every model receives the same brief; divergence emerges in asset generation and animation carryover .

Awwwards Clone: Three Different Tastes From ASI to Minimal Grid

Blind column A opens as an ASI-themed corporate recruiting site : headline ‘ recruiting engineers, quantitative researchers and ML scientists for firms that move markets ’, motto ‘building tools since 2006’ , scattered random logos, a dead search box, SVG-built blocks and staged scroll reveals — messy but with a balanced palette. Column B chooses a portfolio-agency language as Boach Studio / Club Road Coco : filters Art direction / Brand identity / Campaign , cards like Post Code Football Club , heavy SVG placeholders , at times a one-liner that remembers only year/language instead of consent — Simmons notes ‘ would have liked image-generation skills ’ here. Column C breaks the agency rut with Wim Crouwel's grid-based typography experiment Minimal — story ‘ start with a square, program a full alphabet on a minimal grid ’, grids that keep moving as you scroll , a long single-page lesson feel. This third is Simmons' favourite; its length and teaching arc earn bonus points.

Matching references sharpens the table. For the ASI page Grok picked Awwwards Box Studio ‘Visit site’ video — distant resemblance beyond carrying the scroll animation; where assets are missing it was expected to generate them and stays limited to moving blocks. For Aspen Search Fable's output lands very close to the real site — missing some animations but close enough to confuse real vs clone , with block entries and partner strip nearly one-to-one. For Minimal grid Astra's pick is the clear winner : carrying motion and grid from screenshots outclasses the other two SVG fillers. The invoice prices the taste: Fable 5.1 $123 (about 90 million input / 1 million output, ~4,000 seconds), Astra $97, Grok 4.7 $7 (~300 seconds) — the label at the end of build one sticks to Grok: cheap but behind on taste and pixel fidelity .

Keyboard Stage: A 3D That Assembles and Explodes on Scroll

The second build is the keyboard variant of the earlier premium watch page : a single-scroll luxury mechanical keyboard product page with a real-time 3D model at the center — as the page scrolls the model assembles, explodes into parts and is annotated , plus it must respond to typing . The brief is short: premium mechanical keyboard, real-time 3D, everything driven by scroll, typing interaction, full page test . Simmons walks each output with typing and scroll checks ; sound of the key plane, tone and texture enter the notes. The expectation is not 3D alone but how embedded the 3D feels in the web — the scroll animation making it belong to the page.

Column one opens with blue blueprint and explosion phase ; labels overlap at times but the pieces come together as scroll accelerates, spring and gold contact leaf details land in place, with occasional fast-transition glitches and a key shooting through the board . The steel spring and gold leaf texture succeeds, note ‘ 4 mm down, ready ’ and line ‘ Less clutter, more character ’ accompany, with generated typing sound and add-to-cart closing the page. Column two steps ahead with better 3D, smooth scroll and correctly placed labels : interaction ‘ type with your keyboard — type something ’, spec layer Physical Vapor Deposition (not lacquer, does not tarnish or chip) and a more controlled presentation of the key detail ; no part shifts as the model reassembles at the end. Column three tries a different reading with a hand-drawn newspaper texture ; 3D arrives late, sparser and less complex , labels are correct but sparse , the key pops aggressively and the cart component is missing .

Put cost beside load and the table hardens. Fable 5.1: $58 (24 million input, 692,000 output, ~869 seconds) — reference for score and smoothness. Grok 4.7: ~$5 but far longer run and more tokens , output simple 3D and less smooth . Astra: ~$5 and near Fable quality yet far more efficient — hence Simmons' line ‘ Astra, not as good as Fable but incredibly efficient ’. The winner here again is the Fable/Astra line , the loser Grok ; the gap is the model's inability to carry heavy 3D in one shot and to nail scroll physics . The cheap label stays on Grok, quality is shared between Fable and Astra .

Prism Kart and Rainbow Road: A Physics Exam

The third build is a Mario Kart clone ; the map is updated to Rainbow Road and the brief is built around physics . Column A is Prism Cart : flow pick any eraser → start race shows basic illustrations and generic generation , but Rainbow Road-like track and fantastic physics — easy control, boost feel, occasional edge-of-track oddness yet repeatable and playable . Columns B and C's fate is tangled from the start: column B turns left when you press right — the inverted horizontal axis Simmons also saw on the earlier Astra/Fable test — leaving the impression of no QA ; line ‘ I have to press right to go left ’ repeats three times. Column C is entirely scattered : should show rival karts first then spawns at random , unplayable with wayward physics plus the same inverted axis, landing in the worst bucket.

The invoice for once turns against Grok. Fable 5.1: $52, ~3,000 seconds , Astra: $11, ~2,000 seconds , Grok 4.7: $16 — longer run, more input/output tokens for the worst play . Simmons' reveal seals the table: best physics Fable , the inverted-axis bug repeats on Astra , and Grok 4.7 the weakest link . Across three builds it's two clear losses, one losing by a margin for Grok; the cost edge does not cover the revision need because for build three the note stands — ‘no matter how many runs, climbing to Astra/Fable line is hard’ — the price of cheap becomes lost playability .

Stacked together the pattern is blunt: Grok 4.7 trails in heavy code and agent generation in one shot, Fable 5.1 and Astra lead . The second-half recap ties the note to the arena: $2 / $6 versus $10 / $50 — five times on input, eight times on output — and this delta is a real lever in office and knowledge scenarios ; the Grokbot or Grok 4.7 path instead of chat work / co-work closes everyday tasks more cheaply , and the press table holds rows where Grok beats Fable/Astra . By contrast on Terminal-Bench and heavy code Fable 5 (from 26% to 42%) and Astra (around 59%) open a clear gap over Grok's ~26% band ; even if DeepSuite shows 71% with selective framing , the general picture keeps Grok a tier below the frontier . Simmons' note therefore is ‘ a test of heavy code only, limited ’ — knowledge work, automation and daily agent chains are absent from the video.

The close holds two caveats on the same page. First scope : three builds, heavy code, knowledge/automation out of scope ; so Grok's office-automation upside flagged in benchmarks does not appear on this stage. Second the nature of one-shot : no second pass for any model; a 3D page or clone that would recover with iteration is judged by its first output . The economics knot there: at one-fifth the price you can run Grok four to five times for one Fable , yet the Mario Kart rising bill and still distant play hints that even revisions may not close the line . Simmons adds a personal aside: he runs two separate max 20x subscriptions on Chachi and Fable and constantly hits limits ; a serious rival at one-fifth the price is good news for every end user — ‘I just wish it were a bit better’ . The final read is hardware-tinged: xAI owns the compute, so Grok will improve and approach the Fable/OpenAI line ; until then rotate models and experiment is the rule. The comment call mirrors this: what should be built next , leave it below.

Visualization: nodesdaily AI

CursorBench Score (higher is better)

  • Fable 5.151.6%
  • Grok 4.746%
  • Fable 5 (old)26%
  • Astra (missing)—
Grok 4 pts behind at one-fifth the price; Astra omitted from this chart.
TopicStatus
Price lever$2/$6 vs $10/$50; 5× input, 8× output
Awwwards first passGrok $7 cheap but taste/fidelity behind
Heavy build wallOne-shot Grok trails on 3D and Kart
BuildFable 5.1AstraGrok 4.7
Awwwards clone$123 / 4000s$97$7 / 300s
3D keyboard$58 / 869s~$5~$5 long
Mario Kart$52 / 3000s$11 / 2000s$16 long

Key moments

  1. Claim: Fable/Astra parity at one-fifth price
  2. CursorBench 46% vs 51.6% and $2/$6 pricing
  3. DeepSuite 71% and office/legal rows
  4. Method: one bash for three one-shot builds
  5. Awwwards clone prompt and skill
  6. Blind reveal A — ASI recruiting
  7. Blind reveal B — Boach Studio / Coco
  8. Blind reveal C — Minimal Crouwel grid

AI commentary

"Replaying the video line by line I kept one filter: **bill divided by score**. Without putting xAI's CursorBench chart and the three build invoices in the same equation, a $5 vs $58 price tag would tell the story alone; instead I write every paragraph through the lens of **what one shot costs**, testing the ‘five times cheaper’ line against speed, tokens and playability and steering the reader toward the office-automation side where the value holds."

AI assessment

Steel-manning the bull case first: xAI holding $2 / $6 while staying within four points on CursorBench and beating Fable/Astra row by row on office, legal and automation genuinely pulls cost per completed task down; for teams living in Grokbot and multi-hour knowledge work , paying 5× for a few points is not rational everywhere, and the everyday-task speed + price lever makes Grok a sensible pick.

The limit shows on the build rig. The one-shot rule asks Grok to carry heavy 3D and pixel fidelity on the first try, and it drops on taste selection on Awwwards, mature 3D on the keyboard, and playable physics on Mario Kart ; $5 versus $58 on keyboard is cheap but the overlapping labels, glitches and missing cart , the inverted axis + wayward physics on Kart and the SVG filler behind $7 versus $123 on the first build can erase the price edge in practice. xAI owning Cursor and dropping Astra from CursorBench also strengthens the curation read.

On verification the table is testable: $2/$6 versus $10/$50 is consistent on the card, 46% / 51.6% / 71% / 72.7% / 70% are read directly from the chart on video and the cheapest-model tag together with Fable $123 / Astra $97 / Grok $7 and the keyboard/Kart invoices are verified by on-screen counters . Still, ‘best’ is benchmark- and one-shot-sensitive ; Terminal-Bench around 26% and code-heavy work beyond office keeps Grok behind, and a single 71% on DeepSuite does not generalize.

The practical read is selective: for teams with high automation volume, medium context and day-to-day chains on Grokbot/agent harnesses , Grok 4.7 is a strong lever and a price-performance champion candidate; for front-facing work like pixel-perfect Awwwards clones, heavy 3D pages or physics-driven games , Fable 5.1 / Astra remain a tier ahead . Decide on dollars per task and revision need, not on the scoreboard alone ; rotating models and experimenting is the closing advice — leaving the next build idea in the comments is the rational next step.

Sources

6 links; 1 of them also cited by 4 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

grok 4.7 · xai · fable 5.1 · astra · cursorbench · deepsuite · awwwards clone

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

Grok 4.7 in Three Heavy Builds | Nodesdaily