A file smaller than an app icon that beats models ten times its size sounds like marketing, but Needle 3 makes exactly that claim. Its smallest usable slice at 4 layers is about 8 MB, the full 20-layer build is 29 MB. No internet, no server hop, no waiting — the model runs right on the device. The video takes that headline apart piece by piece to see where it holds and where it stretches.
The team behind it is Cactus Compute, a small San Francisco team backed by Y Combinator. Their thesis is blunt: cloud-backed intelligence is powerful but needs a connection, costs money on every call and moves your data elsewhere before answering. They build models and an inference engine to skip that altogether, putting the intelligence on the hardware in your hand or on your wrist. It is not a one-off experiment — they have been shipping open source tools for mobile and edge for a while, and Needle is now on its third generation.
The lineage shows fast iteration. Needle 1 arrived earlier this year as a 26-million-parameter model pitched as a distilled slice of Gemini’s tool-calling ability, tiny enough for a budget phone and fully open from day one. Needle 2 grew to 45 million parameters yet stayed at just 14 MB on disk, using a 256-token sliding window with tool definitions pinned as fixed references, capping memory at around 28 MB no matter how long a conversation ran. The team was clear at the time: it was a specialist, not a generalist — not for long chats or essays, but for picking the right function and filling arguments correctly.
The project’s traction suggests it is not a throwaway release. The GitHub repository sits above 11,300 stars with more than 700 forks and dozens of contributors; the license recently moved from MIT to Apache 2.0 and the commit log shows fixes to the engine and internal state handling landing in the last day or two. Seventeen tagged releases and daily activity point to sustained attention rather than a single drop.
The strongest external signal comes from deployment. Pebble, the company behind smartwatches and the screenless Pebble Index Ring, already runs Cactus Needle inside its app. The founder’s point is simple: a ring with no screen must execute a spoken command every time, with or without connectivity, so they run the model locally instead of depending on the cloud. The footprint stays tiny while performance holds — a real product test beyond benchmark tables.
Needle 3 claims to cover three jobs in one model, all on device. The first is tool calling: give the model the functions your app exposes with descriptions, and it reads what the user says and decides which function or functions to call and which arguments to fill. If a single sentence asks for two things, it returns two calls in the correct order; if nothing matches, it returns an empty list instead of guessing, so the app knows to do nothing rather than do the wrong thing.
The video’s concrete example makes that clear: saying “Dim the bedroom and lock up” to a smart-home app becomes two calls — one to dim the lights, one to lock the doors — executed immediately on device with no round trip to a hub or cloud. That instant offline execution is the exact moment Needle is designed for, especially when voice commands must never be dropped because the network is down.
The second job is structured extraction. You declare the shape you want, hand over messy text, and Needle pulls the exact fields back as typed data. Extracting a vendor name, total amount and due date from an invoice, or key details from a booking confirmation, a phone notification or a form, are the canonical examples. Reliability comes from a decode grammar compiled from the schema that constrains generation so the output is guaranteed to parse and to stay within allowed values. The team notes this approach generalizes to classification, since an enum turns the same mechanism into a labeler.
The third job is text embedding. The same model can turn a sentence into a vector — a numerical fingerprint of meaning — so the app can search, match and route entirely on device. That could be finding the right note among hundreds, picking the closest tool for a vague request, or noticing two alerts say the same thing and merging them. On a watch where screen space and attention are scarce, keeping that logic local is especially valuable; otherwise every query pays a cloud round trip.
Packing three usually separate capabilities into one small binary, if it holds, removes real complexity for builders. What sets Needle 3 apart is what Cactus calls the intelligence ladder: one set of weights where every depth from 2 to 20 layers is itself a complete, usable model, each more capable than the one below. A developer can ship a 4-layer 8 MB slice at 29 million parameters for a watch or the full 20-layer 29 MB build for a flagship phone, all from the same training run. Depths are trained together so smaller subnetworks inherit capability from the full process and each depth can still be fine-tuned independently afterward.
Several architectural bets make that possible. The standard feed-forward block is replaced by a Monarch Hadamard MLP, attention uses grouped-query attention with rotary position embeddings and light causal convolution taps, and the most unusual piece is what Cactus calls an engram — a hash-memory lookup where a large share of parameters lives as memorized patterns rather than live computation, which is why a 121-million-parameter model behaves computationally closer to a 50-million one. Multiple parallel hyper-connections also let information flow across lanes at once. Together Cactus claims these choices cut floating-point operations per token by more than two times versus a standard transformer of the same shape, which maps directly to less battery drain and less heat. The team also notes honestly that sharing weights across sizes is not new in research; what feels new is shipping it this polished for the tiny devices they target.
Before accepting the headline, it helps to see how it was measured. For tool calling Cactus used Mobile Actions at 961 rows of phone commands mapped to Android intents, DroidCall at 200 rows where some requests need two calls in correct order, and Berkeley BFCL v4 at 3,641 rows that also checks withholding a call when no tool applies. For extraction they used DSDC8, Snips Gold and SNIPS 7way, scored as field micro-F1. The comparison setup matters: baselines ran at full 16-bit via a standard engine, DeepSeek V4 Flash was hit via its cloud API, and Needle’s own subnetworks at 20, 16, 8 and 4 layers were run as the real 2-bit quantized binaries that ship to developers. Needle is thus measured in its most compressed, shippable form while baselines are measured at their strongest and uncompressed. The video frames that as a fair “what you would ship today” answer, yet flags that it is not an apples-to-apples compression comparison.
In that frame the headline claim is that Needle 3 outperforms models ten times larger on mobile tool-calling tests and keeps pace with models two to three times larger on extraction. The louder “passes DeepSeek V4 Flash” claim is narrower: fine-tuning on the DroidCall set via the Cactus Platform lifts every subnetwork by 18 to 36 points, and from 4 layers upward the tuned slice passes DeepSeek V4 Flash on DroidCall and Mobile Actions. On Mobile Actions the snapshot is DeepSeek V4 Flash 88.4, Needle 3 full 20-layer 86.0, LFM2.5 1.2B 82.4, Qwen3.5 0.8B 76.0, FunctionGemma 270M 65.1 and Apple’s on-device model 57.6. That is not a claim that a 29-million-parameter model outthinks a frontier model in general; it is a fine-tuned narrow-task win, different from beating a frontier outright but still genuinely useful if your app only needs that narrow task reliably and offline.
The on-ramp is deliberately simple. Install the Python package, fetch the model once from Hugging Face and cache it locally, then decorate a plain Python function — its signature supplies argument types and its docstring becomes the description the model reads. A single run call handles the loop: Needle chooses the appropriate tool, runs your function, incorporates the returned value and sends back a consolidated reply with the execution outcome. When a description alone is not enough, triggers help: regular expressions matched against the user utterance can lock decoding to a specific tool and guarantee a call even below the normal confidence threshold, useful for phrases like “turn the lights on” that must never be missed or misrouted. For extraction the flow is mirrored: declare the shape with a Pydantic model and call extract with messy text and the shape, getting back a typed object with no manual parsing. Every response comes back as one consistent JSON object with function calls, a written-out reasoning trace, a calibrated confidence score and performance numbers like prefill and decode speed plus peak memory. The confidence head has a built-in floor at 0.1 — anything below is moved to a suppressed-calls field instead of executed — and Cactus suggests a three-tier use: above about 0.7 execute immediately, in the middle show and confirm, and when nothing is returned treat it as a refusal because the model found no applicable tool.
Coverage is unusually broad. Cactus ships prebuilt engines each under 1 MB for macOS, five Linux architectures, Windows, three Android architectures, iOS, tvOS, watchOS, a browser WebAssembly build and even a WASI component for specialized embedded targets. That breadth signals an intent to drop into almost any device category, not just phones. Needle 3’s weights sit on Hugging Face and the full source plus engine live on GitHub, free to download and run. The business lives one layer up: the Cactus Platform offers curated data sets, the 2-bit quantization that shrinks the shipped model, evaluation design and tracking, and full-depth fine-tuning infrastructure on Cactus’s own pipeline. In other words the free open model is a strong starting point, and the paid platform is where a team shapes it to its own tools and use case without building a training stack from scratch. That open-core plus paid infrastructure pattern tends to be healthier than a free release with no revenue path — it gives Cactus an incentive to keep improving Needle because revenue depends on builders who want to go deeper, not just download once.
Mobile Actions Scores
- DeepSeek V4 Flash88.4
- Needle 3 20L86.0
- LFM2.5 1.2B82.4
- Qwen3.5 0.8B76.0
- FunctionGemma65.1
- Apple57.6
| Model | Score % |
|---|---|
| DeepSeek V4 Flash | 88.4 |
| Needle 3 20L | 86.0 |
| LFM2.5 1.2B | 82.4 |
| Qwen3.5 0.8B | 76.0 |
| FunctionGemma 270M | 65.1 |
| Apple on-device | 57.6 |
AI commentary
"What strikes me most about Needle 3 is not just size but the way every design choice is measured in battery and heat; the intelligence ladder has lingered in papers for years, and Cactus made it shippable and measurable on a real watch."
AI assessment
Steelmanning the other side, a small specialist beating a large generalist on a narrow task is not surprising — it is the expected outcome. DeepSeek V4 Flash is trained for broad chat, coding, reasoning and long context; on phone-intent suites of 961 rows for Mobile Actions and 200 rows for DroidCall, most of that breadth stays idle. Taking a 29-million-parameter specialist, fine-tuning it on that same narrow distribution for an 18 to 36 point lift and then comparing on the same suite naturally tilts toward the specialist. That does not make the win meaningless; for a product that only needs that narrow task offline, calling a large general model is simply wasteful.
Two limitations deserve weight. First, all Needle scores in the table are measured through the real 2-bit quantized binary, while baselines are measured at full 16-bit. That is the right choice for “what you ship today,” but it is not a compression-equal comparison. An architecture that loses little at 2-bit might gain even more at 16-bit, or the reverse could be true. Second, Needle shines inside its training distribution — consumer device actions for smart home, mobile and wearables, plus structured field filling. Outside that, gains fade: on BFCL v4’s non-Python call surfaces and enterprise API categories, generalization drops, with the gap concentrating in Java and JavaScript SDK slices where the training set has never been.
On verifiability, read the table with balance. The 11,300 stars on GitHub and the live deployment in the Pebble ring are two non-benchmark signals and both point positive, showing maintenance and product trust beyond a chart. Yet benchmark curation itself is company-led, and a company’s own suite will always be kind to its own model. The 115-billion-token pretraining corpus and the details of the 2-bit Cactus Quants pipeline remain proprietary, making exact external replication hard. The extra 18 to 36 points from platform-side fine-tuning raise a fair question: would the same lift appear with the same data outside the platform?
The practical takeaway lands on a clear threshold. If you need a voice command or a short text to become a reliable device action offline — smart home, wearable, robot or microcontroller-class product — Needle 3 is a serious contender, especially with grammar-guaranteed JSON, a calibrated confidence head and trigger regex for must-not-miss phrases. If you need long conversation, open-ended Q&A, creative writing or broad API coverage beyond Python, this is not that tool and Cactus does not pitch it as one. The healthiest decision is not to pick by the 86.0 on Mobile Actions alone but to measure directly with your own tool definitions and your own invoice and notification texts between the 4-layer and 20-layer slices.
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Needle 3 by Cactus Compute
- @cactuscompute.com https://cactuscompute.com/needle
- @huggingface.co https://huggingface.co/Cactus-Compute/needle3
- @github.com https://github.com/cactus-compute/needle
- @gittrend.io https://gittrend.io/repo/cactus-compute/needle
- @cactuscompute.com https://cactuscompute.com/
needle 3 · cactus compute · on-device ai · intelligence ladder · edge ai