On paper the managed-agents bump reads like a routine version note, but in practice it moves the foundation: the old anti-gravity agent ran on Gemini 3.5 Flash, the new build anti-gravity-preview-09-2026 is fully tuned for Gemini 3.8 Flash. The clearest way to feel the gap is the line used in the announcement — before you asked Gemini how to do something, now you hand it the job and walk away. Gemini 3.8 Flash was trained for autonomous, long-horizon work while keeping Flash-grade speed and efficiency, so the shift is less about chatting smarter and more about working through a job end to end.
Every call provisions an isolated Linux computer hosted by Google and the agent moves in. It plans, runs code, browses with Google Search and URL context, creates and edits files, installs packages and loops until the task is complete. The environment is Ubuntu-based with Python 3.12 and Node.js 22 pre-installed, files survive up to seven days of inactivity, and the VM naps after a short idle window then restores on a cold start. The practical implication is different from chat: a single agentic run can consume roughly 100k to 3M tokens, so cost and architecture feel closer to a batch job than to a one-shot response.
Sandbox and Files: Memory Across Sessions
The most useful promise of the sandbox is persistence. A file created in one session is still there in the next, so you can research today and ask tomorrow to turn that research into a seven-day onboarding email sequence without re-uploading anything. The files API reinforces this: upload up to 2 GB per file and let the agent draw from the same dataset across distinct tasks. The walkthrough example makes it concrete — drop 90 days of community engagement data, ask which content types drove the most sign-ups, where the calendar has gaps, and get back a prioritized 30-day content plan, turning a weekly analysis chore into one chain.
A secure bridge to the outside world completes the picture. Credentials are stored outside the workspace and injected at the network layer, never exposed inside the sandbox, so you scope them narrowly before the agent ever sees them. That lets you connect a CRM, email platform or analytics tool and still stay comfortable: plug in the community API, pull activity, filter members inactive for 21 days, draft a re-engagement note per segment and export a ready-to-send outreach sheet with copy included, all in one hands-free retention flow.
Chaining: Three Prompts, One Continuous Project
Multi-step chaining carries both the previous interaction ID and the environment ID forward, so the same workspace keeps building. Step one gathers the most asked questions in the community this month, step two turns them into a structured FAQ document, step three renders that FAQ into a searchable HTML page members can actually use — each step inherits the files from the last. To keep long histories usable, the platform triggers automatic context compaction around 135k tokens, preserving focus without manual pruning. Instead of writing one giant prompt, you hand off a project in parcels the way you would to a human teammate.
For business teams the payoff is immediate and tangible. Hand over a leads spreadsheet, ask the agent to score likelihood to join, segment by interest level and produce an interactive HTML dashboard with the segment breakdown — the output is not advice but a file you can ship. Content research, outreach sequences and retention loops that once ate the better part of a week map cleanly onto the same scaffold. Google's own framing is telling: once the sandbox, infrastructure and execution loop are managed by the platform, teams stop wrestling with orchestration and start productizing the agent's domain behavior.
Not everything is friction-free. Managed agents remain in public preview, schemas can still change, there is no versioning or rollback, and pricing is not chat-like — a single task can burn more tokens than expected, and while environment compute is free during preview the billable shape at general availability is still open. Medium reasoning handles most jobs, high reasoning improves accuracy on dense work at the cost of more tokens, low reasoning saves latency but trims depth. Frameworks such as LangChain, LlamaIndex or CrewAI still matter for complex orchestration, but for teams that want a single API call to stand up a working agent without building a sandbox from scratch, the hosted execution layer is a clear speed win.
AI commentary
"To me this is the first convincing shift from an assistant that talks to one that actually gets work done; once the sandbox and chaining click together, automation stops being a demo and starts shaving real hours off weekly operations."
AI assessment
The walkthrough is compelling as a demo but leans heavily on a single source; nearly every use case circles the same community product, leaving other verticals, failure cases and hard metrics under-explored. That lifts the marketing layer above the technical substance of the upgrade.
What is also missing is clear cost and operational risk framing: how to budget tasks that range from 100k to 3M tokens, how pricing will land after preview, and what happens if a narrowly scoped credential is actually too permissive. The chance that automatic compaction around 135k tokens drops a crucial detail in very long chains, and when to pick Gemini 3.5 Flash versus 3.8 Flash in practice, deserves a fuller discussion.
Directionally, however, the move is right: pushing the hosted execution layer into the platform removes a lot of orchestration toil, especially for smaller teams. The next step should be to step beyond showcase examples and measure, with independent benchmarks, error post-mortems and transparent cost-benefit tables, how many hours these agents truly save in weekly operations.
Sources
7 links; 1 of them also cited by 2 other stories. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube YouTube — Julian Goldie SEO: NEW Gemini Agents Update is CRAZY!
- @ai.google.dev https://ai.google.dev/gemini-api/docs/agents
- @ai.google.dev https://ai.google.dev/gemini-api/docs/antigravity-agent
- @ai.google.dev https://ai.google.dev/gemini-api/docs/managed-agents-quickstart
- @ai.google.dev https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash
- @blog.google https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
Also cited by: The costume on Arena: Gemini 4 traces leaking under the Gemini 3.8 Flash label · DeepSeek V4.1 Flash Meets Gemini 3.8 Flash: A Five-Prompt Coding Duel
- @deepmind.google https://deepmind.google/models/model-cards/gemini-3-8-flash/
gemini 3.8 flash · managed agents · anti-gravity