Krish Naik returns after months without a dedicated paper deep-dive and puts Pathway's Dragon Hatchling (BDH) architecture on the agenda. The thumbnail question is blunt: will BDH replace Transformers? His quick take: BDH is pitched as the missing link between Transformers and brain models, introduced with the September 30, 2025 paper on Hugging Face. The small reasoning model built on top, BDH-CQ, reaches a new cost-efficiency point on ARC-AGI-1 with just 115 to 150 million parameters.
Why Transformers think expensively
Let me unpack why a new architecture is even needed, step by step. When you ask ChatGPT, Claude or Gemini to solve a simple 20-chocolate puzzle — give 7 to a brother, 6 to a sister, how many left? — the model spells out every intermediate step as words: 20 minus 7 is 13, 13 minus 6 is 7. That is Chain-of-Thought, forcing internal computation through language. Like a student who must say every intermediate result out loud, the model spends tokens to think. Tokens mean latency and cost. Pathway calls this the token bottleneck: you can run billions of operations per step inside, but you can only carry a narrow token channel to the next step.
The second bottleneck is memory. A Transformer has two different memories: long-term knowledge baked into frozen weights during training, and a token-indexed KV cache that grows with context and resets each session. Long contexts swell the cache and push systems toward workarounds like summarization or external vector databases. Building an agent in everyday life feels like rewriting the notebook from scratch each time. BDH's thesis is that memory should live on the synapses themselves, not outside; experience should accumulate and change how the next problem is handled. New examples at inference time should directly update internal memory — the difference between an employee on day one and the same person a year later.
How Dragon Hatchling works
BDH takes its cue from the brain's scale-free network. In a brain not every neuron connects to every other; a few highly connected hubs coexist with many sparsely connected nodes in a sparse yet efficient graph. BDH models language as local distributed graph dynamics rather than dense matrix multiplies. Each node talks only to neighbors, memory is carried by edge weights (synapses) rather than node states alone, and the interaction kernel is restricted to sparse spiking-like signals. Reasoning is formalized as an edge-reweighting process. As a toy model, imagine an oscillator network: nodes pull each other's phases and lock into a shared rhythm.
The GPU-friendly variant is BDH-GPU. Here the graph is turned into a tensor-friendly state-space system. Linear attention is aligned to the neuron dimension, activations are kept sparse and positive, and a ReLU-lowrank block replaces the Transformer's MLP. ReLU helps propagate Markov-like signals while preserving modularity. The engineering goal is practical: use LSH to move key vectors into the positive orthant, grow capacity for long context, and naturally obtain monosemantic synapses — each synapse tuned to one concept — alongside sparse activations. Pathway reports both emerging naturally in measurements.
BDH-CQ: learning in context, reasoning in latent space
BDH-CQ is the reasoning specialization of the architecture — the name expands to In-Context Learning with Recurrent Latent Reasoning. The logic is simple: you show a few examples, the model solves the new task only from those demonstrations. It works in two phases. 1) In-context learning: as examples stream in, a recurrent memory is updated, like accumulating short-term notes inside. 2) Recurrent latent reasoning: when the query arrives, the model does not decode its thinking word by word; it iteratively refines the solution in a high-dimensional latent state and decodes only candidate answers. No intermediate step is verbalized, no parameters are updated, no retraining for the task.
The chosen proving ground is ARC-AGI-1. The benchmark hands the system a handful of colored grid input-output pairs, each task with a different hidden rule, and asks it to apply the rule to a new grid. One puzzle might shift a color pattern to the right; the next puzzle needs a completely different transformation. The public evaluation set is where BDH-CQ hits 29.5% pass@2 with a computed inference cost of $0.0007 per task — less than one-tenth of a cent. On the same leaderboard GPT 5.6 Luna (Low) reaches 34.2% but costs about 11 times more, even after OpenAI's July 30 price cut. The result was reproduced in an independent black-box evaluation and in a replication by Lukasz Kaiser, co-author of the 2017 Transformer paper.
Scalability is the part the video emphasizes. Pathway shows early training results up to 600 billion parameters and points to a Transformer-like scaling law. The historical strength of Transformers was predictability: give it more data and parameters and it keeps improving. If BDH follows the same curve, the architectural change offers not just efficiency but a growth path. Naik notes the paper frames this as early evidence, promising but still at the beginning.
What does this mean beyond colored grids? Borrowing Pathway's own analogy, think of planning a complex multi-country trip: flights, visas, budgets, weather, opening hours and preferences all constrain each other. When a flight is cancelled, a capable system should know what stays valid and what must be rebuilt coherently. That is not just retrieval; it needs reasoning across interacting constraints. BDH aims to handle such situations with a continuously updated internal state rather than an external memory patch.
In short, the video converges on one message: the token tax we pay for reasoning is not a law of intelligence but a choice of architecture. BDH tries to turn attention into synaptic memory and weave in-context learning together with latent reasoning into one fabric. We have seen it so far on a limited benchmark with a 150-million-parameter model; the picture will not be complete until it is tested at larger scale, in production integration and on harder reasoning suites.
Key moments
- What is BDH? First look at Dragon Hatchling
Pathway pitches BDH as the missing link between Transformers and brain models.
- The token tax: why every step becomes a word
The 20-chocolate example reveals the hidden cost of Chain-of-Thought
- Memory bottleneck and the external notebook
KV cache swells, agents are forced to write to external vector stores
- Brain-inspired graph: how synapses remember
Scale-free network and reasoning as edge reweighting
- BDH-CQ: learn in context, reason in latent space
Examples update recurrent memory, solution refines in latent state
- 29.5% on ARC-AGI-1 at $0.0007
A new efficiency frontier at less than one-tenth of a cent per task
AI commentary
"For me the real claim of BDH is not the percentage point but how it cuts the bill for thinking: not having to turn every reasoning step into a token could finally make reasoning scalable in production."
AI assessment
Steel-manning the BDH thesis, it looks like this: the cleanest way to cut the cost of reasoning in language models is not a longer chain-of-thought but an architecture that can carry thought without turning it into words. Remove the token bottleneck and both latency and bill drop, letting you quietly test more candidate solutions within the same budget. Scoring 29.5% on ARC-AGI-1 — where the rule changes every task — for less than one-tenth of a cent per task lowers the unit price of intelligence, and the independent replication by Lukasz Kaiser adds credibility.
The limits are equally clear. First, BDH-CQ is a small model specialized for colored grid puzzles; the same efficiency has not yet been shown on general reasoning, math, coding or long-document tasks. Second, the signal up to 600 billion is based on early training observations, not a full-scale model above 150 million going head-to-head with frontier models on the leaderboard. Third, the Hugging Face paper and the BDH-CQ report were released at different stages; operational metrics like latency, memory footprint and error-recovery behavior in production are not yet public.
On verifiability the picture is fairly transparent: the official score on the ARC Prize leaderboard is shown next to a computed hardware-time cost per task, with the independent black-box test report linked. But comparison costs for some models reflect API pricing, so it is not strictly apples-to-apples; and ARC-AGI-1 itself is heavily visual-abstraction focused and does not one-to-one represent real-world agent tasks. A stronger verdict would require the same protocol on ARC-AGI-2 or live agent benchmarks.
The practical takeaway: if you need to move reasoning into production where it must be billable, the BDH approach is a compelling experimental lane. For high-volume, low-cost reasoning — continuous rule extraction, small agent loops or research on a constrained budget — trying a 150-million-parameter model makes sense. For safety-critical or high-stakes financial decisions, there is not yet evidence that it replaces general-purpose frontier models; the architectural bet still has to prove itself with scale.
Sources
7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Krish Naik: Will BDH Replace Transformers?
- @pathway.com https://pathway.com/research/bdh-explainer/brain-inspired-ai-architecture
- @pathway.com https://pathway.com/introducing-bdh-cq
- @pathway.com https://pathway.com/blog/pathway-150m-model-breaks-arc-agi-1-cost-efficiency-frontier
- @huggingface.co https://huggingface.co/papers/2509.26507
- @arxiv.org https://arxiv.org/html/2608.09888
- @arxiv.org https://arxiv.org/html/2509.26507v1
bdh · dragon hatchling · pathway · arc-agi · post-transformer