722 manuscripts and 372 result families
The week's headline fits in one sentence: OpenAI probed roughly 4,000 mathematics research problems with an unnamed frontier model and released the resulting 722 manuscripts as 372 result families . The host treats the collection as the first major preview of the unreleased model; the Bel name circulating online remains unconfirmed. The scale alone gives pause: a single model, in a single move, produced as much text as a mathematics department spreads over years.
The first detailed account came from ScientificAmerican: according to the magazine, a substantial share of the results has passed Lean verification , meaning the logical skeleton of the proofs is machine-checked. The wave landed right after OpenAI's Navier-Stokes result a month earlier, an atmosphere the magazine sums up as hundreds more results raining onto a field already in shock. OpenAI also consulted circles around the Institute for Advanced Study and independent mathematics and AI advisory groups on how to share the results.
NewScientist supplies the frame for what the scale means: the outlet recalls the trajectory from models struggling with high-school mathematics in 2019 to frontier mathematics within a few years. This time the emphasis is on breadth rather than depth, and the single-day release of hundreds of findings provoked both admiration and bafflement across academia. According to NewScientist, mathematicians will need months to determine whether the proofs carry genuinely new ideas or repackage existing techniques. That cautious mood should accompany the rest of the week's reading.
The host's structural point cuts deeper: the story is not a single theorem but research itself becoming parallelizable . Instead of one mathematician probing one direction for months, thousands of research directions can be scanned at once. The chain is familiar: mathematics feeds physics, physics feeds engineering, and engineering feeds chips, energy, materials, and medicine. As AI helps discover better algorithms, architectures, and hardware, that help returns as better AI. This feedback loop explains why the number 722 is more than a curiosity.
Day two of the 28-day shipping marathon
The mathematics wave coincided with day two of OpenAI's pledge to ship something new every day for 28 days. According to AIReport's rundown, the day's package holds four items: auto-review is now free for all ChatGPT users and no longer consumes plan quota; developer tiers shrink from five to three (build, launch, grow), with the top tier's lifetime payment requirement cut from $1,000 to $500. A quota-burning feature going free is a direct cost reduction for anyone running long-lived agent work.
The other half of the package sits in meetings and decision infrastructure. A new meetings plugin for ChatGPT takes notes during calls, produces personalized summaries with next steps, and saves everything into ChatGPT's workspace; combined with the memory feature, those notes can flow into project plans and coding agents. The fourth item is the Luna-powered Decisions API entering public beta. Per AIReport, the marathon's rule is harsh: any missed shipping day triggers a full usage reset for everyone.
Tiered cyber access and Europe's open-weight push
On the Anthropic front, the week's news is the expansion of the Cyber Verification Program. According to Anthropic's announcement, the program now has three access tiers: a defense tier covering incident response, malware analysis, and vulnerability research; a red-team tier opened to authorized penetration testing; and a specialized tier granting broader capabilities to organizations passing extensive security review. Opus 5.5, Sonnet 5.5, and Mythos 5.1 are reachable through these tiers, while generally available models continue under the same safeguards. Applications are open to all qualifying teams.
From Europe came the week's boldest open-weight announcement: Mistral put Large 4 (nickname Le Chonk) into public preview. Per Mistral's technical note, the model carries 1 trillion parameters yet activates only 49 billion per step thanks to its mixture-of-experts design ; it was developed end to end in Europe and served from the company's own European cloud. Handling text, code, images, and audio in one architecture, its weights are due at the end of October while the preview API is usable today. For European companies, data staying on the continent is the announcement's standout aspect.
VentureBeat's behind-the-scenes account also prices the scale: the model trained from scratch over roughly two months on 4,000 accelerators in the company's own data centers and covers more than 160 languages. The company claims the best results among open models on enterprise workloads such as cybersecurity and manufacturing, and beyond closed frontier models on visual grounding . The sharpest exhibit is a reverse-engineering trial: the model identified a never-before-seen malware sample as Cobalt Strike and produced a full investigation report with configuration dump, indicators of compromise, and a Yara rule; per VentureBeat, work that normally fills a researcher's day finished in 12 minutes.
The host's own WoAI Bench setup adds a caveat: Large 4 beats GPT-6 Luna on domain tasks such as software engineering and agentic tests, yet trails the same model on speed and efficiency. The verdict is blunt: slowed by its reasoning process, this model is not a daily driver but a candidate for heavy domain work such as cybersecurity and knowledge tasks. A natively multimodal build and a 1-million-token context window support that verdict.
The Google front and the week's software notes
On Google's side the image-generation front is moving: Nano Banana 2.1 was announced as an upgrade surpassing its predecessors across the board. DeepMind's model card places Nano Banana 2.1 in the Gemini 3 family; it jointly understands text and image inputs, produces image and text outputs, and supports a 1-million-token context window. As the quality race in image generation accelerates, this release's largest gains cluster in the visual domain; the card's measurement details live on DeepMind's page.
The quiet but significant embedding move is EmbeddingGemma 2: only 740 million parameters, mapping text, code, images, audio, and video into one shared space, built to run on phones and computers. According to Digit, the model opened its weights under the Apache 2.0 license. A real-world trial shared by the host backs the claim: 300 animal photos across 30 species were indexed on DGX Spark and 200 similarity queries run against 20 separate images; the model scored 165 correct matches, 82.5 percent precision, perfect 10-out-of-10 on zebras and pandas while struggling with horses and leopards.
Two notes from the software front. First, Claude Code 2.1.292: per the release-note roundup by PromTime, an effort parameter on the agent tool lets users set how hard each sub- agent thinks, while local protocol servers now speak the 2026-07-28 protocol by default. A free-to-access stealth model inside OpenCode is also making the rounds; the host saves the details for a later video. Second, from the open-source world: a developer announced recreating seven major Adobe applications, Photoshop included, as free alternatives written in Rust; Photocraft, the Photoshop counterpart, reportedly reaches 90 percent of the original's capabilities in its first release. The project's GitHub repository is public, though the developer concedes these builds are not yet ready to replace Adobe in daily professional use.
| Development | Why it matters |
|---|---|
| 722 math manuscripts | A frontier model serializes research |
| Mistral Large 4 in preview | 1T parameters, weights in October |
| Claude Code sub-agent dial | Separate thinking depth per sub-agent |
Key moments
AI commentary
"The pattern this week is unmistakable: what scales is no longer just the models but the volume of research they produce. The 722 manuscripts remain a promise until independently verified, yet the size of the promise already reveals the field's tempo. My real question is whether Mistral Large 4's weights will genuinely go public."
AI assessment
The strongest objection concerns the numbers' interpretation: the 722 manuscripts have not yet passed independent review; Lean checking validates a proof's logic, not its significance. Mathematicians face months of sifting before anyone knows how much of this text carries genuinely new ideas and how much repackages existing techniques. Breadth is not evidence of depth.
The list of unknowns is long: the identity, training, and access schedule of the Bel-nicknamed model; Large 4's pricing and license terms until the weights land; the Luna-backed Decisions API still in beta; and no word on who stands behind the stealth model. These gaps counsel reading the week's headlines as an opening lap, not a victory lap.
The host's position deserves a note too: the video's spine is his own WoAI Bench harness, and the broadcast carries a Skool community pitch plus a Zapier SDK sponsorship. That does not make the selection biased, but it is worth knowing: which tests get the spotlight is related to the tools in the tester's hands.
The practical takeaway compresses into three lines: security practitioners can review Anthropic's application page; teams hosting data in Europe can watch Large 4's weight calendar; meeting-heavy readers can try ChatGPT's meetings plugin. For the rest, one rule holds: a single verified breakthrough out of 722 papers outweighs hundreds of unverified announcements.
Sources
11 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — WorldofAI
- @scientificamerican.com Scientific American
Also cited by: Nasdaq on top: nuclear deal, memory rally and the $40B AI race
- @newscientist.com New Scientist
- @anthropic.com Anthropic
- @mistral.ai Mistral
- @venturebeat.com VentureBeat
- @deepmind.google Google DeepMind
- @digit.in Digit
- @aireport.net AIReport
- @promtime.net PromTime
- @github.com GitHub — Photocraft
artificial intelligence · openai · anthropic · mistral · google · open weights · mathematics