Back to feed

10,000 Agents, 88 Hours, $1 Million: AI Mastermind #39 From Code to Cash to Autonomy

From OpenAI's claim of 10,000 agents cracking Navier-Stokes in 88 hours to Tesla Cybercab's pedal-free brake, from Astra beating Sol on low effort to Stripe Link's wallet for agents, AI Mastermind #39 packs 20+ threads; I mapped every thread to see what is hype and what is immediately useful.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — i7IYtC3v_38
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

AI Mastermind episode 39 opens with Jirka Herník, Míra and Michal in one room and one question: have you let your bots pay yet? It is not a joke; with Stripe's Link wallet and the Machine Payments Protocol a single sentence can buy placement, and the crew shows how they bypass a European block with a Revolut virtual card.

10,000 Agents, 88 Hours: The Million-Dollar Equation

The headline claim comes from OpenAI: around 10,000 concurrent agents around an unreleased model solved the Navier-Stokes existence and smoothness problem in about 88 hours. The crew walks through the 165-page draft and Lean formalization said to have landed on September 5; the Clay Institute still lists the problem as unsolved and OpenAI says it will not claim the prize.

Scale is larger than the headline number. The Navier-Stokes run is described as 2.7 million messages and 130 billion output tokens, with 4.9 million messages and 300 billion tokens across the broader attempt. Codex is cast as the consolidator that merges intermediate results and fans them out to agent groups, with compute cost talked about in the millions of dollars but without a breakdown.

My note: 88 hours is catchy, but the denominator matters — what hardware and parallelism. The model is unreleased, the orchestration is closed, and the bundle that matters for mathematicians — the Lean repo and logs — is the real test; the rest needs peer review before it becomes history.

Pedal-Free Tesla: Cybercab and the Brake Question

The Tesla Cybercab scene has no wheel and no pedals. No master cylinder, no fluid, no lines; each caliper has its own electronic actuator. The dry brake-by-wire described by InsideEVs and The Drive is one of the first production uses after Chery's Exeed EX7, and steering is also steer-by-wire with the box behind the front motor.

The regulatory gamble is discussed openly: could the car be blocked for missing pedals, and did Tesla price that in while building dozens anyway? The stated engineering upside is lower build cost and per-wheel force control; suppliers named are Brembo, ZF and Bosch.

From Europe the picture lags. In London, Waymo and Uber robotaxi plans wait on DVSA permit and TfL operating consent; a continental test that was penciled for late 2026 may slip into 2027. The studio joke lands: America drives, we are still debating permits.

Why Europe Hits the Brakes: Regulation and Safety

Safety is framed with an old image: a boy walking in front of early cars with a flag. Regulation came after the cars, and autonomy will be similar. Even if the robotaxi crashes ten, a hundred or a million times less than a human, it will still crash; the question shifts from zero crashes to governance.

In practice the issue is not fluid but accountability. Brake- and steer-by-wire centralize trust in software; the crew argues for transparent logs, versioned software and staged city approvals rather than a promise of flawlessness.

Model Wars: Astra, Fable, Sol and the Second Best

A flurry of models in the last 14 days confuses names: GPT-6 Astra, Fable, Sol, Terra, Luna. Tibo's calibration line sums it up: Astra on low effort beats Sol on high. OpenAI's Terminal-Bench Science has Astra at 64.6% vs Fable at 52.6% with lower cost, and Agents Last Exam at 59.3% ahead of Opus 5 at 55.5%.

The studio draws a practical line: you no longer need the biggest model for every job. Where Sol was once ignored in favor of Terra, now Astra low or medium and even Luna high handle sub-agents. It mirrors phones and PCs: chasing the flagship every time stopped paying off, the second best now covers most days.

Subscription economics heat up. Claude Code is described as eating limits, with weekly usage at 95% and a 150 limit ending next week; on the Codex side three resets in four days and unprecedented demand for Astra are reported. One host cancelled Claude and moved to Astra only, others warn to move quickly before new signups are paused.

How Will Agents Pay? Stripe Link and llms.txt

Stripe's move is crisp: on September 8 it announced Link's wallet for agents integrated with Muse. At over one million Link-accepting businesses checkout is instant; elsewhere Link issues a single-use virtual card for the approved amount. The consumer approves the total in chat, the agent never sees the card, and the flow rides on the Machine Payments Protocol with stablecoins or cards.

Jirka's live demo grounds the theory: he adds llms.txt to Výsluní.cz, tells any agent — Claude, Codex, Grok — kup první pozici and it buys placement. The file tells agents how to use the site and how to pay; the design is polished between Claude Design and Codex with a reusable magic prompt.

The European block appears again: agent payments are restricted, so the workaround is a Revolut virtual card handed to the agent. The crew's take, half joking: if you grant payment permission the agent spends, if not the flow stops; the bottleneck is consent design, not capability.

New Toys: Muse, Your Clone and the Game Factory

Meta's Muse arrives the same week: a personal agent with a secure cloud environment, payments via Stripe and approval flows, with a free tier plus $20 and $100 plans. Privacy is pitched as the main selling point — the studio smiles at the pitch from Meta but takes the architecture seriously: no spend without approval, no card sharing.

On the creative side Codex writes the behind-the-scenes. The episode's edit, micro-animations and zooms were built with Codex; the crew notes DaVinci Resolve 21.1's MCP support via Python and an agent that models in Blender and imports into a game engine to build a game from one sentence.

The cloning joke has substance. After Michal's episode 38 cloning gag, the crew now clones with visuals — an agent rebuilding a participant from existing recordings. The demo is short, the implication long: when image, audio and code share orchestration, production time collapses to hours.

Do Benchmarks Lie, or Does the Harness Win?

The benchmark curtain is pulled hard. After refreshes, Gemini falls from 89% to 19% and Muse from 88% to 33%, while GPT-6 Fable and Astra hold their scores; Deep SWE and Terminal Bench examples show how fragile old scores were. Takeaway: freshness of the test matters as much as the score.

The second act is sharper: same model, same task, same time — different harness, different outcome. CivBench runs 300+ turns with 76 MCP tools and measures Proactive Monitoring Rate and RAG@10; Claude Code's harness leads, KimiCode, OpenCode, Hermes and Exo lag. The message: the harness is part of performance.

UX comparison follows. Claude Code feels visually polished with live sub-agent monitoring and easy switching; Codex's /agents gives a static list with little interaction. Speed and small copy-paste quirks shape daily comfort; one host admits missing Claude Code's feel.

herdr: Herding Agents With One Rust Binary

The practical star is herdr. One Rust binary, no Electron; inside Warp it multiplexes agents as real terminals. States are simple: blocked, working, done; instead of hunting panes, herdr tells you who needs input, and both keyboard and mouse are first-class.

Setup is one line: curl -fsSL herdr.dev/install.sh | sh. Server, client and update split cleanly; the same layout lives on laptop, VPS or rented box. It does not wrap SSH, it bridges it; agents use a socket API to split panes, read output and wait on each other. A plugin marketplace extends it.

The trick is survival. Herdr's server lives in the background, remembers every session ID; even after a reboot the layout returns and each agent hears continue. One narrator says he delayed reboot for three days before herdr; after herdr he rebooted and returned to finished agents. Detach is ctrl+b q, reattach is herdr.

Stepping back, episode 39 cannot be reduced to one headline but the thread is clear: scale grew, the second-best model now suffices, harness choice moves the needle and a simple multiplexer like herdr makes living with autonomous agents practical. My call is to leave big claims to peer review and build the small efficient steps today.

Visualization: nodesdaily AI

Key moments

  1. Intro: Have you let your bots pay?We tell the agent in one sentence to buy first placement.
  2. 10,000 agents, 88 hours and a million-dollar claim
  3. Tesla Cybercab: No wheel, no pedals, no fluidEach caliper has its own electronic actuator.
  4. Astra on low beats Sol on highAstra low outperforms Sol high.
  5. Stripe Link: Wallet for agents and virtual cards
  6. herdr herd: Close, return, continueOne Rust binary to manage the whole herd from one place.

AI commentary

"I did not binge this episode in one pass; I split my notebook in three, because a 16,500-word conversation shifts axis every five minutes and skipping a detail makes the next joke fall flat."

AI assessment

Steel man: treat OpenAI's 88-hour, 10,000-agent Navier-Stokes claim as a demonstration until independently checked. Clay still lists the problem as open, the Lean repo awaits community review, the model and orchestration are closed; 130 billion tokens and millions in compute suggest scale, not correctness, without an itemized cost and reproduction.

Limitations: dry brake-by-wire removes hydraulic redundancy; single-point electronic failure and fallback layers get little time in the episode. Benchmark resets shatter old scores but CivBench itself is a 23-run pilot, too small for a model ranking; task selection and playbook matter as much as the harness gap.

Incentives and checkability: the claimant sells the unreleased system it showcases. Tesla's cost saving trades against regulatory risk, Muse's Stripe tie grows both wallets. The check is not the press release but the open log and a reproducible Lean build.

Practical take: my line is to build small efficiency now and let big theory wait for review. Use Astra low or Terra/Luna high for most work to protect monthly limits, require single-use cards and in-chat approval for agent payments, and keep terminals alive with herdr; the rest can wait until peer review lands.

Sources

9 links; 2 of them also cited by 18 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

ai mastermind · openai · navier-stokes · tesla cybercab · muse · herdr · stripe

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…