Alibaba opens with a sharp question: how many taps does it take to book a trip on your phone? The claim behind Qwen Intelligence, unveiled on September 22 at the Apsara Conference in Hangzhou, is that the number should be zero. This is not a chatbot you download or a Qwen update that pops up in your browser; it is a full agentic stack handed to phone makers to bake agents directly into the handset. Honor is the first partner, co-building the next generation of AI phones, with the Honor Magic9 series and the Honor Robot Phone slated as the first devices on September 28. In the words of Steven Hoi, head of Qwen Mobile AI at Token Foundry and VP of Alibaba Token Hub, agentic smartphones are meant to be the primary gateway between personal AI and the physical world — the phone stops being a device you operate and becomes a device that operates for you.
What Qwen Intelligence Is — and Is Not
The architecture is three layers, each solving a different bottleneck. At the bottom sits a Qwen foundation model tuned for mobile scenarios from the outset, not a huge general model squeezed into a phone, so it runs with lower on-device cost. In the middle lies the modular platform layer: independent model deployment, harness customization, unified tool governance and automated evaluation, plus device-cloud coordination that shifts work to the phone or the cloud by speed and cost. On top are premium agents packaged for vertical needs. A maker can take the whole stack or pick modules and wire them into its own system. That modularity, alongside Alibaba’s disclosed RMB 380 billion AI and cloud infrastructure plan, frames Qwen as an operating system of the AI era rather than a single model release.
Three Agents, Three Jobs: Planner Thinks, Use Acts, Creative Makes
The launch trio makes the division of labor explicit: the Mobile Planner Agent thinks, the Mobile Use Agent acts, and the Mobile Creative Agent makes. The Planner is the brain — planning, task decomposition, tool orchestration and dynamic adjustment. Alibaba’s go-to demo is arranging a business trip to Shanghai for Monday, which it splits into booking flights, reserving a hotel, creating calendar events and planning the itinerary. If a flight cancels, memory and proactive services rebook and suggest alternatives without asking you to re-explain the whole trip; the context is already held. The Creative agent is deliberately light, built for on-phone image generation and editing that turns a single sentence into a usable creation in about 3 seconds, claimed as roughly twice as fast as the nearest rivals.
What keeps the Planner honest is the harness — a layer that holds memory and state, preserves constraints across long turns, and aligns what you said with what tools actually returned. Ask it to change one setting and it edits just that, without rewriting everything. The demos highlight behaviors where most assistants stumble: catching a mismatch between conversation and a tool’s report and trusting the tool, waiting for a download to finish before naming the file in the next app, and recovering from a failed conversion by leaning on the calculator instead of quitting. On benchmarks, the Qwen Planner Agent 27B posts 77.05% overall on MobilePA-Bench, first among the evaluated systems and 9.83 points above its own baseline — yet the gap to second place remains narrow, so leadership is measurable but not a blowout.
Use Agent: API First, GUI as Fallback
The clever bit in the Mobile Use Agent is not just where it taps but how it chooses to act. It follows an API-first, GUI-fallback hybrid: if an app offers a proper structured connection it uses it because it is faster and more reliable; if not, it falls back to operating the screen like a human taps, swipes and navigates. Alibaba stresses strict security boundaries and privacy protection while doing so. The disclosed numbers are strong: 82.1 on MobileWorld, 92.2 on MobileWorld-Real and 97.5 on AndroidDaily, with a reported 90% end-to-end success rate. Gadget Pilipinas cites up to 91.8% task accuracy on Honor AI phones on flows exceeding 100 steps. The same scorecard puts ByteDance, OpenAI, Anthropic and Google models behind on those mobile benchmarks. Method matters more than the headline score: training builds on 100+ smartphones and 150+ applications for task creation, trajectory collection, training and evaluation on real handsets, not simulators.
That real-device emphasis is critical because the wild is where agents usually break. MobileWorld-Real was built for it: more than 400 tasks across 100+ apps. The report’s rationale is blunt — sandboxes cannot reproduce apps that change layout overnight, permission prompts, captchas, surprise pop-ups or a network drop halfway through step 40. The Use Agent’s move set reflects that reality. It is not tap-only; it can click, long press, type, open apps, drag, hit back or home, run bash commands, call APIs directly and ask you a question mid-task. Better still, it does not move one-by-one — it completes several moves in a single decision, part of why it feels faster. The footprint is not phone-only either: the same agent covers computer and browser use, logging 79.5 on OSWorld-Verified and 73.6 on WebArena, with roughly 40% of its desktop tasks involving batched CLI commands.
Three-Second Images and Four Open Benchmarks
The Creative agent’s lightness is a deliberate trade: getting from one sentence to a usable image quickly, without pushing every job to the cloud. Alibaba frames 3 seconds as about two times faster than leading competitors, but the more under-appreciated move is opening the benchmarks themselves. Four are public: MobilePA-Bench for planning, MobileWorld for cross-app execution, MobileWorld-Real for real-device performance and MobileWorld-Safety for safety. MobilePA-Bench is not small: 1,705 evaluation tasks, 212 real-world tools, 13 functional domains and 89 subcategories, testing tool use, memory, skill use and sub-agent collaboration. MobileWorld is Apache-licensed on GitHub, with 201 tasks across 20 mobile apps, and its paper is heading to ACL 2026.
How MobileWorld is built explains why its results are trusted. It runs in a Docker container with a rooted Android virtual device inside. Instead of pointing at mutable live services, it hosts its own open-source apps: Mattermost for team chat, Mastodon for social and Mole for Uni for shopping. Because the backend is theirs, they can check the database directly to confirm a message was truly sent. A device snapshot is taken, so every run starts from an identical phone — same starting point, reproducible results. Two task types fill gaps most benchmarks skip: tasks where the agent must come back and ask you something, and tasks that use Model Context Protocol tools, so the test is hybrid, not just screen taps. That design also explains why device-cloud coordination sits at the center of the platform, not as an afterthought.
Who Gets It and When
Who can touch it today depends on which lane you are in. If you are a general user, the honest path is buying a phone that ships with it; the Honor Magic9 series and Robot Phone are the first. If you are a developer, Qwen Intelligence is available via its official website and Alibaba Cloud services, and the benchmarks are fully open — you can download MobileWorld, run it in Docker and test your own agent on the same tasks. If you are a researcher or just curious, technical reports for both the Planner and UI agents are public, with recorded demos that expose tool calls, reasoning and where failures and recoveries happen — a rare level of transparency. On the roadmap are personalized memory with proactive assistance, multimodal interaction and vertical agents from ecosystem partners. Memory, when done right, is the threshold that turns a tool into an assistant that actually helps, and Alibaba flags it as still in the pipeline, meaning everything shifts again once it lands.
The line most viewers will miss is how selection works. These agents search apps, compare options and pick one thing for the user; you never see a list of ten results, you see one answer. The same shift is already happening in Google AI Overviews, ChatGPT, Perplexity and Gemini: fewer results means more decisions made for the user. So the question is not whether your content ranks, but whether an AI picks you when someone asks. Qwen Intelligence is built in three layers for that reality — a mobile-tuned foundation at the base, a platform in the middle that governs tools and auto-evaluates, and packaged agents on top. Makers can take the whole stack or select pieces, much like layering an operating system rather than bolting on a feature.
What to Watch Next: Memory, Speed and the Selection Economy
Three takeaways are worth carrying. First, stop framing this as phone-only; the Use Agent already handles computer and browser work, so the discipline of writing clear step-by-step instructions for an agent pays off everywhere — practice it now. Second, study the API-first, GUI-fallback pattern as a lesson for your own workflows: give your AI a direct connection when you have it; screen control is backup, not default, because it is slower and more brittle. Third, watch the memory layer; personalized memory with proactive assistance is cited as in the works, and when it lands the assistant crosses from reactive tool to anticipatory helper. The near-term test is the Honor launch on September 28 and whether the open benchmarks reproduce: not the leaderboard, but repeats on 100+ real handsets will tell us the real story.
Key moments
- Zero taps idea — why a trip should take zero taps
- What Qwen Intelligence is and is not — platform not app
- Three agents introduced — Planner thinks, Use acts, Creative makes
- One-answer economy — not ten results but one choice
- Planner harness and memory — rebook after cancellation
- Real-device test — 100 phones 150 apps 400 tasks
AI commentary
"My read: Alibaba is not trying to win the agent race with bigger models but by moving into the phone itself. The three-agent bundle feels less like a single assistant and more like an OS layer — memory, tool governance and device-cloud balancing in one kit. That makes the Honor move more than a hardware deal; it is a live showcase for whether the stack holds up in the wild."
AI assessment
Strength: the work moves evaluation out of simulators onto real handsets at scale — 100+ phones and 150+ apps with trajectory collection — and the headline numbers (77.05% on MobilePA-Bench and 82.1/92.2/97.5 for the Use agent) read as a layered proof, not a single score. Four benchmarks opened, with MobileWorld Apache-licensed, is unusually transparent about reproducibility.
Limits and open questions: leadership is thin — the Planner’s gap to second place is small and the 90% end-to-end claim depends on coverage; the app mix and network/permission scenarios where it dips are buried in detail, not the headline. Strict security and privacy boundaries are asserted but stop thresholds for payments, data deletion and permission asks are not shown with audit logs. The 3-second image claim is single-prompt; how much remains in heavier multi-step edits is unclear.
Skeptic’s take: the one-answer economy is as much risk as convenience — as agents pick for you, invisible choices multiply and any biased or sponsored pick vanishes inside a single answer. And API-first is only as strong as apps’ willingness to open an API; at the most critical moments the agent can fall back to the most brittle mode, leaving brand promise and lived experience apart.
Practical takeaway: for phone owners the Honor launch after September 28 is the real test; for builders, running MobileWorld in Docker on the same tasks is a cheap, honest mirror; for teams, give the agent a direct path where an API exists and keep screen control as fallback until the memory pipeline lands. Repeats and trace literacy, more than the scoreboard, will decide what holds.
Sources
8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Alibaba Qwen Intelligence: Three Mobile Agents
- @gadgetpilipinas.net https://www.gadgetpilipinas.net/2026/09/alibaba-qwen-intelligence-agentic-phones/
- @alibabacloud.com https://www.alibabacloud.com/en/press-room/alibabacloudunveilsstrategicroadmaps?_p_lc=1
- @digitalphablet.com https://digitalphablet.com/ai/alibaba-launches-qwen-ui-agent-surpassing-gpt-5-6-and-claude-4-8/
- @honor.com https://www.honor.com/global/news/honor-magic8-china-launch/
- @pandaily.com https://pandaily.com/alibaba-qwen-intelligence-phone-agent-stack-honor-magic9
- @news.cocoloop.cn https://news.cocoloop.cn/en/2026/08/qwen-ui-agent-real-device/
- @techadvisor.com https://www.techadvisor.com/article/2249860/app-less-ai-phone-terrible-idea.html
alibaba qwen intelligence · honor magic9 · mobile agents · apsara 2025 · mobileworld