AI Mastermind episode 39 opens with Jirka Herník, Míra and Michal in one room and one question: have you let your bots pay yet? It is not a joke; with Stripe's Link wallet and the Machine Payments Protocol a single sentence can buy placement, and the crew shows how they bypass a European block with a Revolut virtual card.
10,000 Agents, 88 Hours: The Million-Dollar Equation
The headline claim comes from OpenAI: around 10,000 concurrent agents around an unreleased model solved the Navier-Stokes existence and smoothness problem in about 88 hours. The crew walks through the 165-page draft and Lean formalization said to have landed on September 5; the Clay Institute still lists the problem as unsolved and OpenAI says it will not claim the prize.
Scale is larger than the headline number. The Navier-Stokes run is described as 2.7 million messages and 130 billion output tokens, with 4.9 million messages and 300 billion tokens across the broader attempt. Codex is cast as the consolidator that merges intermediate results and fans them out to agent groups, with compute cost talked about in the millions of dollars but without a breakdown.
My note: 88 hours is catchy, but the denominator matters — what hardware and parallelism. The model is unreleased, the orchestration is closed, and the bundle that matters for mathematicians — the Lean repo and logs — is the real test; the rest needs peer review before it becomes history.
Pedal-Free Tesla: Cybercab and the Brake Question
The Tesla Cybercab scene has no wheel and no pedals. No master cylinder, no fluid, no lines; each caliper has its own electronic actuator. The dry brake-by-wire described by InsideEVs and The Drive is one of the first production uses after Chery's Exeed EX7, and steering is also steer-by-wire with the box behind the front motor.
The regulatory gamble is discussed openly: could the car be blocked for missing pedals, and did Tesla price that in while building dozens anyway? The stated engineering upside is lower build cost and per-wheel force control; suppliers named are Brembo, ZF and Bosch.
From Europe the picture lags. In London, Waymo and Uber robotaxi plans wait on DVSA permit and TfL operating consent; a continental test that was penciled for late 2026 may slip into 2027. The studio joke lands: America drives, we are still debating permits.
Why Europe Hits the Brakes: Regulation and Safety
Safety is framed with an old image: a boy walking in front of early cars with a flag. Regulation came after the cars, and autonomy will be similar. Even if the robotaxi crashes ten, a hundred or a million times less than a human, it will still crash; the question shifts from zero crashes to governance.
In practice the issue is not fluid but accountability. Brake- and steer-by-wire centralize trust in software; the crew argues for transparent logs, versioned software and staged city approvals rather than a promise of flawlessness.
Model Wars: Astra, Fable, Sol and the Second Best
A flurry of models in the last 14 days confuses names: GPT-6 Astra, Fable, Sol, Terra, Luna. Tibo's calibration line sums it up: Astra on low effort beats Sol on high. OpenAI's Terminal-Bench Science has Astra at 64.6% vs Fable at 52.6% with lower cost, and Agents Last Exam at 59.3% ahead of Opus 5 at 55.5%.
The studio draws a practical line: you no longer need the biggest model for every job. Where Sol was once ignored in favor of Terra, now Astra low or medium and even Luna high handle sub-agents. It mirrors phones and PCs: chasing the flagship every time stopped paying off, the second best now covers most days.
Subscription economics heat up. Claude Code is described as eating limits, with weekly usage at 95% and a 150 limit ending next week; on the Codex side three resets in four days and unprecedented demand for Astra are reported. One host cancelled Claude and moved to Astra only, others warn to move quickly before new signups are paused.
How Will Agents Pay? Stripe Link and llms.txt
Stripe's move is crisp: on September 8 it announced Link's wallet for agents integrated with Muse. At over one million Link-accepting businesses checkout is instant; elsewhere Link issues a single-use virtual card for the approved amount. The consumer approves the total in chat, the agent never sees the card, and the flow rides on the Machine Payments Protocol with stablecoins or cards.
Jirka's live demo grounds the theory: he adds llms.txt to Výsluní.cz, tells any agent — Claude, Codex, Grok — kup první pozici and it buys placement. The file tells agents how to use the site and how to pay; the design is polished between Claude Design and Codex with a reusable magic prompt.
The European block appears again: agent payments are restricted, so the workaround is a Revolut virtual card handed to the agent. The crew's take, half joking: if you grant payment permission the agent spends, if not the flow stops; the bottleneck is consent design, not capability.
New Toys: Muse, Your Clone and the Game Factory
Meta's Muse arrives the same week: a personal agent with a secure cloud environment, payments via Stripe and approval flows, with a free tier plus $20 and $100 plans. Privacy is pitched as the main selling point — the studio smiles at the pitch from Meta but takes the architecture seriously: no spend without approval, no card sharing.
On the creative side Codex writes the behind-the-scenes. The episode's edit, micro-animations and zooms were built with Codex; the crew notes DaVinci Resolve 21.1's MCP support via Python and an agent that models in Blender and imports into a game engine to build a game from one sentence.
The cloning joke has substance. After Michal's episode 38 cloning gag, the crew now clones with visuals — an agent rebuilding a participant from existing recordings. The demo is short, the implication long: when image, audio and code share orchestration, production time collapses to hours.
Do Benchmarks Lie, or Does the Harness Win?
The benchmark curtain is pulled hard. After refreshes, Gemini falls from 89% to 19% and Muse from 88% to 33%, while GPT-6 Fable and Astra hold their scores; Deep SWE and Terminal Bench examples show how fragile old scores were. Takeaway: freshness of the test matters as much as the score.
The second act is sharper: same model, same task, same time — different harness, different outcome. CivBench runs 300+ turns with 76 MCP tools and measures Proactive Monitoring Rate and RAG@10; Claude Code's harness leads, KimiCode, OpenCode, Hermes and Exo lag. The message: the harness is part of performance.
UX comparison follows. Claude Code feels visually polished with live sub-agent monitoring and easy switching; Codex's /agents gives a static list with little interaction. Speed and small copy-paste quirks shape daily comfort; one host admits missing Claude Code's feel.
herdr: Herding Agents With One Rust Binary
The practical star is herdr. One Rust binary, no Electron; inside Warp it multiplexes agents as real terminals. States are simple: blocked, working, done; instead of hunting panes, herdr tells you who needs input, and both keyboard and mouse are first-class.
Setup is one line: curl -fsSL herdr.dev/install.sh | sh. Server, client and update split cleanly; the same layout lives on laptop, VPS or rented box. It does not wrap SSH, it bridges it; agents use a socket API to split panes, read output and wait on each other. A plugin marketplace extends it.
The trick is survival. Herdr's server lives in the background, remembers every session ID; even after a reboot the layout returns and each agent hears continue. One narrator says he delayed reboot for three days before herdr; after herdr he rebooted and returned to finished agents. Detach is ctrl+b q, reattach is herdr.
Stepping back, episode 39 cannot be reduced to one headline but the thread is clear: scale grew, the second-best model now suffices, harness choice moves the needle and a simple multiplexer like herdr makes living with autonomous agents practical. My call is to leave big claims to peer review and build the small efficient steps today.
Key moments
- Intro: Have you let your bots pay?
We tell the agent in one sentence to buy first placement.
- 10,000 agents, 88 hours and a million-dollar claim
- Tesla Cybercab: No wheel, no pedals, no fluid
Each caliper has its own electronic actuator.
- Astra on low beats Sol on high
Astra low outperforms Sol high.
- Stripe Link: Wallet for agents and virtual cards
- herdr herd: Close, return, continue
One Rust binary to manage the whole herd from one place.
AI commentary
"I did not binge this episode in one pass; I split my notebook in three, because a 16,500-word conversation shifts axis every five minutes and skipping a detail makes the next joke fall flat."
AI assessment
Steel man: treat OpenAI's 88-hour, 10,000-agent Navier-Stokes claim as a demonstration until independently checked. Clay still lists the problem as open, the Lean repo awaits community review, the model and orchestration are closed; 130 billion tokens and millions in compute suggest scale, not correctness, without an itemized cost and reproduction.
Limitations: dry brake-by-wire removes hydraulic redundancy; single-point electronic failure and fallback layers get little time in the episode. Benchmark resets shatter old scores but CivBench itself is a 23-run pilot, too small for a model ranking; task selection and playbook matter as much as the harness gap.
Incentives and checkability: the claimant sells the unreleased system it showcases. Tesla's cost saving trades against regulatory risk, Muse's Stripe tie grows both wallets. The check is not the press release but the open log and a reproducible Lean build.
Practical take: my line is to build small efficiency now and let big theory wait for review. Use Astra low or Terra/Luna high for most work to protect monthly limits, require single-use cards and in-chat approval for agent payments, and keep terminals alive with herdr; the rest can wait until peer review lands.
Sources
9 links; 2 of them also cited by 18 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — AI Mastermind Episode 39
- @runtimewire.com https://runtimewire.com/article/openai-10000-ai-agents-navier-stokes-proof
- @insideevs.com https://insideevs.com/news/807121/tesla-cybercab-brakes-no-fluid/
- @stripe.com https://stripe.com/newsroom/news/stripe-helps-meta-muse-shop-with-link
Also cited by: Meta Muse: 5 Surprising Features Worth Trying
- @herdr.dev https://herdr.dev/docs/quick-start
- @github.com https://github.com/herdrdev/herdr
- @openai.com https://openai.com/index/gpt-6-astra/
Also cited by: Brain and Body: A Single-Screen Agent Setup with GPT-6 Astra on Hermes · Space Bunny Alpha: Inside OpenRouter's Free Anonymous AI Experiment · Robot-Use Agents: Why General-Purpose Models May Win Robot Control · Building a Productive Card Collection App in Minutes with Base44 and GPT-6 Astra · Price War Begins: GPT-6 Sol and Luna Halve Model Costs · Gemini 4 Leak? 10 Interactive 3D Tests Against GPT-6 Astra and Fable 5.1 · Gemini 4 Pro Leaks, GPT-6 Soul in Testing: From Arena to Google Cloud, the Week's AI Shockwave · Cloning a Channel With One Prompt: The $33K Video Factory Built on GPT-6 Astra and Higgsfield · GPT-6 Astra Guide: How Horizontal Power Turns the Model Into Work Done · From Hand Sketch to Realistic Villa: A Showcase Video with GPT-6 Astra and Higgsfield MCP · The $500-a-Day Claim With GPT-6 Astra: Building Three Business Models End to End · When Agents Take the Job: Investing in the Stack After GPT-6 Astra (+5)
- @arxiv.org https://arxiv.org/abs/2609.02459
- @independent.co.uk https://www.independent.co.uk/travel/news-and-advice/driverless-taxis-uber-wayve-waymo-tfl-b3040914.html
ai mastermind · openai · navier-stokes · tesla cybercab · muse · herdr · stripe