Models released quietly under preview codenames, then forcing a full retest at official launch: the host is openly tired of this cycle. Yet with viewers insistently writing LongCat in the comments, he rolled up his sleeves. LongCat 2.0 had been so bad it never even made his coding leaderboard and faded from memory. Now 2.5 is out and free to try for two weeks; the conditions look right for a short video and a quick verdict.
The test routine has two layers, as always. First 24 coding prompts are run and passing tests are counted; then the quality of the generated code is scored under GPT-5.6 judging. The quality review looks at two separate tracks: a React and TypeScript frontend inside a single project, and a PHP and Laravel backend on a separate track. The ceiling is 60 points and the host performs the update live on screen; expectations start low because the model comes from neither an established lab nor a glossy track record.
Its Place on the Scoreboard
When the table refreshes, the prediction is confirmed: LongCat 2.5 Preview appears in the lower ranks, not rock bottom but far from the summit. Scoring 44 points in total, the model never approaches leaders hovering between the 50s and 57. The most striking detail the host puts on record is the language split: just 0.9 out of 5 on the Go project against 4 out of 5 on Dart and Flutter. Four out of five projects show zero failed tests, so the model finishes most jobs — but the time it takes to cross the finish line tries one's patience.
This is the portrait of a model that finishes yet lags. The run saw intermittent failures and response times wore the host down; were it not free, it would never make a recommendation list. Still, the verdict is not entirely dark: it was judged sufficient for small tasks and wrote its name on the board with a nonzero score. There is no leap closing the gap to the top, but after the nonentity performance of 2.0, 2.5 at least shows a measurable presence.
Price, Identity, and the Free Queue
The model's identity card deserves attention. LongCat is the AI arm of food-delivery giant Meituan; according to LongCat's official blog, version 2.0 carries a 1.6-trillion-parameter MoE architecture, with about 48 billion parameters active per token and 1 million tokens of context. Per Reuters, the model was trained on a 50,000-chip cluster built entirely from domestic processors. In-house measurements shared in the official GitHub repository read 70.8 on Terminal-Bench 2.1 and 59.5 on SWE-bench Pro, and the open-weights page on HuggingFace confirms the picture. On the official site the 2.5 preview price matches 2.0 — 30 cents input and $1.20 output per million tokens — with a discount label attached, openly flagging that the price may rise later. One telling detail: the model is absent from OpenRouter, so access flows in practice through OpenCode.
The free-access side is the video's most practical section. Per OpenCode's Zen documentation, the model identified as longcat-2.5-preview-free is free on both input and output, and the catalog rotates other zero-price options such as MiMo-V2.6-Flash Free and MiMo-V2.5 Free. In the TUI, the list opened with the /models command brings these models forward instantly with the free filter. For price comparison the host cites the GPT-5.6 Luna band: 20 cents in, $1.20 out. LongCat costs a little more at 30 cents input while GPT-6 Luna is far cheaper; yet in the host's view LongCat's quality trails GPT-5.6 Luna and runs slower, so cheapness alone never settles the choice.
The MiMo Fix and the Preview Dilemma
The video's surprise subplot is the previous day's topic: MiMo-V2.6-Flash. An hour after that video aired, the model's team announced they had diagnosed and fixed a tool-call repetition bug expected to raise speed and possibly quality. The host ran a single-query check but caught no visible improvement; the model still felt very slow, with a full retest planned for later in the week. Meanwhile MiniMax waits in the queue: per MiniMax, the open-weight M3 scores 83.5 on BrowseComp, while the MiniMax-M3 card on HuggingFace reads 80.5 on SWE-bench Verified and 59.0 on SWE-bench Pro. In the finale the host lays an honest dilemma on the table: test the preview now or wait for the official release, especially while theories circulate that official releases sometimes get quietly weakened? He leaves the answer to the comments and invites anyone with deeper LongCat knowledge to enrich the debate.
Key moments
AI commentary
"An independent coding test treating a preview model with this much distance is valuable; the scores say more than polished launch tables. LongCat 2.5's free window is worth trying, but it is too early for production code."
AI assessment
The strongest counterargument is this: a single independent test, however transparent, does not decide a model's fate. The host's 24-prompt routine feels more authentic than in-house benchmarks because it rests on real user tasks, yet a single run over a small sample is open to statistical noise. Moreover, the judge's chair is occupied by the GPT-5.6 family, and judging by OpenCode's Zen list, members of that same family are also competing in the race — the referee and the contestant are relatives. So read the 44 points not as an absolute verdict but as one lab's single measurement.
There are gaps in the video too. The sense of slowness is never quantified; without tokens-per-second, queue times, or retry rates, the speed complaint stays at the level of impression. The root cause of the 0.9-out-of-5 collapse on the Go project is never investigated; whether the model does not know the language or the test harness is fragile remains unclear. No link is drawn between the slowness and the 50,000-chip domestic cluster claim reported by Reuters. And the host's theory about post-preview score drops hangs in the air as an untested suspicion.
The host's possible interest deserves a note as well. A format centered on free models drives steady traffic toward the rotating free models on OpenCode's Zen list, and every new preview means a new video. But publishing the scores uncensored and writing the model into the lower ranks shows this format is no advertising brochure. Placing LongCat's official-blog records like 70.8 side by side with the video's 44 points gives the reader a healthy balance.
The practical takeaway for readers is clear. For small scripts, trial projects, and satisfying curiosity, the longcat-2.5-preview-free queue is a sensible stop; you learn the model's character without spending a penny. But for work with deadlines, for languages like Go, and for agent loops where speed matters, stay loyal to the top models. Before turning preview scores into a buying decision, wait for the official release and at least one independent retest.
Sources
8 links; 2 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — LongCat 2.5 Preview coding test
- @opencode.ai OpenCode Zen — models and free pricing
Also cited by: Gemini 4 Pro Leaks, Rogue Agents and China's 10-Trillion-Parameter Plan: One Day of AI Headlines
- @longcat.ai LongCat official blog — LongCat 2.0 specs
- @github.com GitHub — meituan-longcat LongCat-2.0 repository
Also cited by: Gemini 4 Pro Leaks, Rogue Agents and China's 10-Trillion-Parameter Plan: One Day of AI Headlines
- @reuters.com Reuters — Meituan domestic-chip training report
- @huggingface.co Hugging Face — LongCat-2.0 open weights
- @minimax.io MiniMax — M3 model page
- @huggingface.co Hugging Face — MiniMax-M3 model card
longcat · artificial intelligence · coding · opencode · mimo · minimax