A million tokens means several thick books held together, and this model promises that in a pocket-sized package. The free 4-billion-parameter edition ships as a compressed four-bit file of only 2.6 GB. The question starts right there: can an ordinary computer with no graphics card actually run it? The short answer splits in two. The model opens, yes; but using the full million in daily work, no.
The model's identity clarifies the picture. This openly licensed release from iFLYTEK's Spark team, together with its 1.7-billion-parameter sibling, is among the rare designs with a native million-token window on device. It covers more than 200 languages and uses an efficiency-first hybrid attention architecture: alternating sliding-window layers with full-attention layers bring long-text cost down. The details of this design are documented on the model card at huggingface.co.
The first condition for card-less use is a current runtime. Right after release the model failed on standard llama.cpp builds with an unknown-architecture error, and support landed officially on September 6. The card lists minimum versions separately for llama.cpp, Ollama and LM Studio, and the maker publishes the model on Ollama under its own account. File choice matters too: the default tag pulls the 8.2 GB full file, while the sensible pick for modest memory is the 2.6 GB four-bit build. I confirmed this split on the model page at ollama.com.
There are two different speeds here and they must not be confused. Read speed is how fast the machine digests your text, write speed is how fast it generates the reply. On an 8-core Xeon box writing crawled at about 6 tokens a second while reading landed between 30 and 36. On an 18-core ARM machine writing reached 47 and reading scaled up to 254. The notable pattern is that extra workers aided intake yet harmed output on that ARM silicon: generation crested with six workers and sank under 7 tokens a second once sixteen workers contended. Every one of these figures comes from short prompts and says nothing about giant-document behavior. There is an easy setup trap because the product advertises a million-token ceiling. The first tester found the runtime reaching for the full million by default and had to clamp the context size by hand. Write the number yourself rather than leaving it to the default. That silent defaults cause headaches is also confirmed by the Ollama context guide on neuralgist.com, where the client can pick a small window based on memory conditions.
To grasp what a massive prompt costs in memory, look at the small records the model stores while reading. This KV cache entry per token is consulted while processing everything that follows. By the arithmetic, a million tokens need about 39 GB for the cache alone, slightly over 41 GB together with the model itself. So a desktop with 64 GB of RAM comfortably holds the whole document. A 32 GB machine squeezes in only with a compressed cache whose accuracy price is unmeasured, and a 16 GB machine hits the wall around a couple of hundred thousand tokens.
The reason the bill stays manageable is hidden in the design. If all 36 layers stored the whole text, the arithmetic points at roughly 155 GB. Instead 27 layers keep only the latest 512 tokens, and only nine full-attention layers save the cache across the entire text. llama.cpp uses that saving by default. At standard 16-bit precision the cost per token falls to 36 KB: about 1.2 GB at 32K, 4.8 GB at 128K and just under 39 GB at the full million. The matching figures on a community GGUF card and the 1.3 GB reserve measured at 128K with four-bit compression back this up; the GGUF page on huggingface.co gives the file and architecture detail.
Even with enough memory, the real bill arrives as waiting time. Nobody has published a stopwatch test of a long prompt on a processor for this model, so every long duration here is arithmetic worked out from measured short-prompt read speeds. On the 18-core ARM setup a 32K prompt takes about 3 minutes, 128K lands around 20 minutes and the full million stretches to roughly 13 hours on that same fast machine. On the 8-core Xeon box with four threads it balloons to about 4 days. The reason is quadratic attention cost : in the nine full layers each incoming token is compared against every earlier token in the cache. By the math the comparison equals all other work at 51K tokens and costs about 20 times more at a million. The background of this in-memory behavior is explained by the llama.cpp cache-architecture notes on deepwiki.com. The gap to even a small graphics card is striking. In the GGUF author's measurement a little card read a short prompt at 1684 tokens a second and the full 131K tokens at 557, yet finished the job in about 4 minutes. That identical load means a twenty-minute pause on the quick ARM processor and roughly two-point-five hours on the eight-core server unit. Even the quickest processor intake mark reported anywhere, from Apple M4 Max efficiency cores inside another software stack at 345 tokens per second, still implies around nine and a half hours per million. So where RAM suffices, the million-token window is mostly a waiting problem; tens of thousands of tokens are what a processor can actually read in minutes.
Even if you accept the wait, a second question remains: do answers hold up at that length? On the maker's own table the model scores 56.3 on the ArtificialAnalysis long-context reasoning test, which reasons over roughly 100K-token documents per question, against 57.0 for its closest 4-billion rival Qwen3.5-4B. Two caveats are mandatory: the number comes from the vendor grading its own model rather than an independent lab, and 100K is one tenth of a million. Above that line there are no published retrieval scores, reasoning tests or needle-in-haystack results. What the test measures is explained by the long-context leaderboard on artificialanalysis.ai.
Before you ever reach a million there is a daily processor trap waiting: thinking mode. The model writes hidden working notes before the final answer by default. At the 8-core machine's pace every thousand tokens of inner monologue costs nearly 3 minutes. In a coding task from the HuggingFace discussions the model spent the entire 800-token budget across 233 seconds inside its own head and delivered zero lines of code. In another shopping total it stalled at the 4,096-token limit, yet solved the same problem cleanly in 549 tokens with the mode off. A 10,000-token loop debating one parameter name at 128K context would mean roughly half an hour of waiting on that Xeon box. Turn thinking off for quick jobs and keep it only for deep reasoning with patience to spare.
So is a card-free PC able to operate million-token AI? The hardware passes, yet daily use of the full window fails. It truly launches on an ordinary processor, and in theory a 64 GB desktop retains a million tokens, but the decisive barrier is delay rather than capacity. Even on the faster tested machine the arithmetic says about 13 hours of uninterrupted fan noise just to read the prompt, with no published proof that answers hold up past about 100K. The one big exception is acceleration: a small card read about 131K tokens in 4 minutes where the faster processor needed around 20. On a processor, grab the official 2.6 GB four-bit file on a current runtime, dial the window down to tens of thousands of tokens and turn thinking off for quick jobs. The industry context for this launch is given by the iFlytek report on yangtzeer.com, presenting it as China's first open million-token edge model.
Key moments
AI commentary
"The million-token headline is exciting, but after this test the number stuck in my head is waiting time. That the model runs at all is impressive, yet the full context on a processor is a patience project, not a daily driver."
AI assessment
The strongest counterpoint is that slow does not mean worthless. A model that runs offline and keeps documents on the machine is invaluable for privacy-sensitive work. Someone queuing batch summaries overnight can accept a 13-hour read as a fair price. For phones, robots and field PCs with no graphics card, having a small open model that runs agents locally is a win in itself.
The gaps are real. Nobody has timed a million-token prompt on a processor, so every long duration here is arithmetic derived from short-prompt speeds. There is no independent accuracy evidence above roughly 100K tokens, no published retrieval or needle-in-haystack scores. The accuracy cost of a compressed cache is unmeasured, and the claim that a 32 GB machine can hold a million tokens rests on that unknown option. Trusting the full window before these gaps close is speculation.
A note on the speaker's position: The Stack did not run the model itself; the measurements belong to testers on the HuggingFace page and to the independent tester Javier Canete. The long-prompt times, memory tables and the 51K crossover point are the speaker's own arithmetic. That transparency is good because we can tell measured numbers apart from calculated ones. Still, arithmetic may understate the slowdowns real hardware will show.
The practical takeaway is clear. On a card-less machine, take the official 2.6 GB four-bit file on a current runtime, keep the window at tens of thousands of tokens and switch thinking off for quick jobs. If feeding library-scale documents is routine, even a small graphics card takes you from hours to minutes. Read the million-token claim as a roadmap, not as today's working setting; the llama.cpp memory-management notes on deepwiki.com explain well why this hybrid cache design behaves this way.
Sources
7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — The Stack: million-token CPU test
- @huggingface.co HuggingFace: XHToken Spark-X2.5-4B model card
- @ollama.com Ollama: SparkLLM Spark-X2.5-4B
- @artificialanalysis.ai ArtificialAnalysis: Long Context Reasoning Benchmark
- @deepwiki.com DeepWiki: llama.cpp memory management and KV cache
- @yangtzeer.com Yangtzeer: iFlytek million-token edge AI model
- @neuralgist.com NeuralGist: Ollama context window guide
artificial intelligence · local model · long context · cpu inference · memory · open source