In this video Fatih Bellisoy walks through an AI agent he presents as usable without any subscription or API key, built around Z.ai's GLM 5.3. The model is framed as a computer-use system for desktop tasks, long-horizon planning and the design of complex workflows. The ability to experiment without cost pressure is the reason he now uses it in his own projects.
GLM 5.3 keeps the same base as GLM 5.2; every lift comes from scaling post-training. The team combined IndexShare for long-context efficiency, SAO for RL on long horizons and slime for large-scale asynchronous training, then grew the set of task environments collected over months. The Flash variant starts from a freshly trained base: 320 billion total parameters with only 18 billion active, a hybrid of sparse and linear attention plus Manifold-Constrained Hyper-Connections that cuts attention compute by 3x and KV cache by 4.4x at 1M context; on Z.ai Code Bench Flash jumps from 46.2 on 5.2 to 63.4, closing in on Claude Opus 4.8.
The first demo is intentionally rough: a single short prompt for a 2+1 suite hotel room in 3D that you can walk through. The result is not perfect, yet a navigable space from one sentence validates the idea of pulling visual feedback into the coding loop.
The second project is a personal, still-unfinished control center that orchestrates several agents and surfaces live work data — social metrics, view counts and latest videos in one place. It is incomplete but has been pushed forward with GLM 5.3 at no extra cost. The dashboard shows the agent can read desktop context beyond just writing code.
The third demo is a 3D chess game, again from a single prompt. Play proceeds cleanly, turns alternate and animations run without breakage. The knight model is weak on aesthetics, but shipping a rule-correct, interactive game from a terse command shows the model can handle front-end and game logic together.
You can also try the model on the web. Z.ai's chat surface offers GLM 5.3 and a faster Flash option with templates for slides and web apps. On desktop the normal route is API access with usage-based billing: Z.ai lists GLM 5.3 at $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.26, while the Flash variant drops to $0.15 in and $0.50 out and $0.03 cached. Flash hits 57 on the Artificial Analysis Intelligence Index at $0.045 per task, delivering similar intelligence at roughly one-tenth the price. Freebuff claims to zero out that bill with ad support.
After downloading Freebuff you pick an empty folder; the agent gains read-write access to that workspace. In the model picker GLM 5.3 Flash appears as recommended. The service grants 25 tokens per day, 5 tokens buys a one-hour GLM 5.3 session. Alternatives are listed too — Mimo 2.5 at 10 tokens per hour, DeepSeek V4 Flash at 25 — and you can also attach Claude Code or Codex via your own key. The session does not start on app launch but on your first message, kicking a one-hour countdown.
For the final build planning mode is enabled to create a chat interface on a canvas where answers can be pinned as notes. The agent asks scoping questions, drafts a detailed plan and waits for approval before writing code. After revision the app lets you select a reply and add it to the canvas, which the host finds handy for language learning. Including setup and small fixes the build took about 15 minutes.
AI commentary
"What struck me most was not the price but the sense of freedom — the hourly session starts on first message, not on app launch, and the plan needs your approval before any code runs."
AI assessment
The strongest counter to this video, as I see it, is that a flashy one-prompt 3D scene is not the same as production code. Without measures of test coverage, maintainability or total cost over time, free can look better than it is. Z.ai's own table shows GLM 5.3 at 88.2 on Terminal Bench 2.1, yet on later-stage exploitation benchmarks it still trails GPT-5.6 Sol and Mythos 5; scaling post-training does not lift every rung equally.
Limits the video does not mention matter in practice. Z.ai delayed open weights by about two weeks because cyber capability grew faster than expected — ExploitBench more than doubled from 24.4 to 54.4 and ExploitGym jumped from 29 to 105 tasks in two hours, forcing extra hardening. On Freebuff, 25 tokens a day means roughly five hours, full mode is available in 25+ countries and elsewhere you fall back to limited mode behind VPNs, and while text ads fund the free tier, transparency around data handling and traffic routing remains thin.
On incentives and verifiability I keep two numbers apart. Holding GLM 5.3 API pricing flat at $1.40 in / $4.40 out sounds stable, yet the real bill moves with context window, caching and whether you are routed to Flash at $0.15 / $0.50 and $0.03 cached — different labels for different products. Freebuff's ad-funded routing has not faced independent audit; for a desktop agent with folder access, network traces and model-selection policy deserve a second-source check before you grant it a sensitive repo.
In my own use I place this stack as ideal for students, solo developers and tight-budget teams who want fast prototypes, 3D experiments or a language-learning board via GLM 5.3 Flash on Freebuff; shipping a canvas-pinned chat in about 15 minutes is a real speed gain. For enterprise work that needs compliance, audit trails or strict privacy over long contexts, waiting until weights are open and the ad model is clearer is the more cautious call.
Sources
8 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — GLM 5.3 + Freebuff Video
- @z.ai https://z.ai/blog/glm-5.3
- @z.ai https://z.ai/blog/glm-5.3-flash
Also cited by: AI Tier List Reset: GPT-6 Astra Takes the Crown as Subscription Math Rewrites the Ranks
- @freebuff.com https://freebuff.com/
- @github.com https://github.com/CodebuffAI/freebuff
- @huggingface.co https://huggingface.co/zai-org/GLM-5.3-Flash
- @curiouslm.com https://curiouslm.com/blog/glm-5-3-open-weights-cyber-delay
- @aiinsiders.net https://aiinsiders.net/article/zai-keeps-glm-53s-api-price-flat-the-real-bill-isnt
glm 5.3 · freebuff · z.ai · ai agent · desktop agent · free ai · coding agent