Anyone who has flipped on the memory switch in a chatbot and expected it to remember everything has hit the same wall: the next day, in a fresh chat, you cannot pick up where you left off. This video argues the auto-generated memory is not broken but deliberately kept thin, and builds a three-step ladder: account-wide memory, project memory, and your own file-based memory system.
Account-wide memory is the set of facts and preferences the chatbot writes about you on its own, and every note there flows into each new chat. The demo in the video is blunt: a micro iPhone pitch worked on for hours, with the Satya Nadella and Tim Cook detail, confirmed first-version slides and a visual polish step waiting next, is gone the next morning. Opening the memory page under personalization shows no trace of the project, only high-level lines about role and writing taste.
The reason is a design choice: anything stored at the account level gets dragged into every future chat, and one wrong entry poisons everything. The many-hats example carries the point: a manager at work, a parent at home, plus spouse, sibling, side business, weekend golf; what helps in one context is noise in another. If each chat had to carry unrelated context, output would get worse, so the robot chooses to remember less.
Level one has two patches and both are half-measures: telling the robot mid-chat to update its memory works, you can watch the memory page change, but remembering to ask every time stays your job, and the new entry lands in the same thin global profile and leaks into unrelated chats. Pulling the latest meeting notes through connected tools like Google Drive helps only so far; if that work never had a meeting, or the knowledge sits in an earlier session inside the chatbot, the connector stays blind. The Claude and Gemini versions behave the same.
Level two draws a boundary with the projects feature around a single work stream, and that boundary changes everything: because every chat inside the project discusses the same work, the robot can keep far more specific notes without smearing them across other work. The account-level layer keeps cascading underneath, so each project chat receives both the global profile and the project context. Project knowledge is now retrieval-backed too: according to the vendor help docs, once project content grows, the system fetches only the relevant slice, pushing capacity to multiples of the old limit.
The live demo shows both faces of that boundary: asking the Claude project for the next steps of the pitch returns the right answer, version one done and visual polish next. But asking for the full list of confirmed attendees returns a partial list, even though a screenshot of the calendar invite was pasted into a chat a week earlier and read correctly at the time. The robot simply did not judge those names worth keeping. Imagine a stale date leaking into a status note for your manager, says the author; a fair fear.
The most important sentence of the video lands here: in everything so far, the author is still the robot. The robot still holds the pen: it picks which details survive, which folder they sit in, and the moment they are stored. The patch is familiar: telling project memory the correction directly, adding the new attendee by name, fixes the list. But noticing the problem, remembering the fix, and telling the robot stays human work; while memory was supposed to lift that load, the load stays with us, and the arrangement looks unsustainable.
Level three changes the envelope: memory stops being a black box and becomes plain-text files you can open, read, and edit. You write the rules, the robot does the upkeep. All three major labs now ship system products for this: Claude Cowork and Claude Code on the Anthropic side, ChatGPT Work and ChatGPT Codex on the OpenAI side, Gemini Spark on the Gemini side. The heart of the design is one small routing file: at the start of each task the robot reads it and routes the request to the owning project folder, so even with more than twenty active projects only the relevant one gets loaded.
In the live demo a fresh session with no context is asked where the micro iPhone pitch stands; the system first reads the root memory file, finds the pitch among active projects, and opens the project files inside the consulting folder for a full read. The result impresses: the deck is content-complete, a visual polish round waits, the date moved from September 15th to October 6th, three new attendees force a larger meeting room, and open items are listed. The critical point is that none of this came from a black box: the date sits in a plain-text file open in Obsidian, and you can change it to October 30th by hand.
The closing protocol is the real charm of the system: a single wrap-up command ends the work session and the robot scans the session for decisions, learnings, and progress worth keeping. It separates what needs approval from what it handles itself: turning a one-off acronym fix into a permanent rule is asked as a proposal, while logging the headline edits and updating the root memory file happen automatically. The headline pass from minutes earlier proves it: acronyms in slide titles get cleaned by one plain-language remark within about a minute, and the rerun confirms the titles are acronym-free.
The finale weighs the trade honestly: level one for high-level lines that should follow you across chats, level two for those buried in a single work stream, level three for those carrying more than twenty projects, writing rules, and manager updates in one workspace. Control grows as you climb, but so do the setup week and a work discipline that differs from the chat window; where to stop is a call to make against that trade.
AI commentary
"Watching this video made me question my own setup: I have carried context to chatbots in project folders for months, and I have lived the missing-attendee answer many times. The third rung tempts me, but its price deserves honesty, because an unkept file is worse than a forgetful robot. So instead of echoing the demo, I put the objections from independent research on the table."
AI assessment
Let me steelman against the video first: a memory that remembers everything is not an unqualified good. Reporting on 2026 research from the AI company Writer, TechCrunch describes how memory systems pull models toward user misconceptions; once a favorite book is stored as Station Eleven, the model names it even for unrelated questions, and the tilt grows stronger with compression tools like Mem0 and Zep. The second finding bites harder: fed finance misconceptions, a model that judges a company correctly with memory off fails with memory on. So each rung of the ladder may amplify the same risk; more context does not always mean better judgment.
There is a second mechanical limit the video never tests, sitting ironically inside its own demo: in retrieval-backed projects the robot does not read everything, it fetches only the slice it deems relevant. The March 2026 vendor help doc states this openly; capacity rises to multiples of the old limit but fetching stays selective. The attendee-list case may therefore be architecture, not taste: if the names never entered the fetched slice, they never made the list. The twenty-project routing file at scale also goes untested: colliding project names, moved folders, rotting files, and conflicts invisible in a one-person demo get no hearing.
Two separate alarms ring on verifiability. A January 2026 MIT Technology Review piece argues agents collapse once-separated contexts such as work, health, budgeting, and private correspondence into single unstructured pools, and the stored detail can seep into shared pools once the agent connects to outside apps; the same piece notes the move to pipe personal history into Gemini through Personal Intelligence. On security, the attack class reported by CSO describes hidden instructions planted in robot memory with a single prompt. Both qualify the comfort of files sitting on my disk: even self-hosted files get read aloud by an agent that talks to connected apps with the same mouth.
My verdict: level two suffices for anyone buried in a single work stream, where the project boundary plus a correction discipline carries the load. But for someone running many parallel efforts with house writing rules and regular status notes, the file-based system is the real answer; its price is a setup week and a discipline unlike the chat window. The cost the video undersells is this: the system wants a gardener, and an unkept memory file is more dangerous than a forgetful robot, because this time the wrong answer arrives at full stride.
Sources
7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube Jeff Su — episode video (YouTube)
- @techcrunch.com https://techcrunch.com/2026/06/10/how-memory-tools-can-make-ai-models-worse/
- @technologyreview.com https://www.technologyreview.com/2026/01/28/1131835/what-ai-remembers-about-you-is-privacys-next-frontier/
- @simonwillison.net https://simonwillison.net/2025/May/21/chatgpt-new-memory/
- @support.claude.com https://support.claude.com/en/articles/11473015-retrieval-augmented-generation-rag-for-projects
- @csoonline.com https://www.csoonline.com/article/4213632/new-attack-lets-hackers-plant-hidden-instructions-in-ai-memory-with-
- @zdnet.com https://www.zdnet.com/article/i-tried-chatgpts-memory-function-and-found-it-intriguing-but-limited/
ai memory · chatgpt · claude · productivity · knowledge management