The video opens as a sequel to a short memory video published four months earlier. That video had offered a quick fix for Claude Code reading unnecessary files and burning tokens, but its short length left most viewer questions unanswered. The new video aims to close that gap, arriving just as agent-memory debates flare up again. The author sets the frame upfront: the whole explanation runs on Claude Code because it is both widely used and a concrete example whose file structure and tracking tools make the topic tangible.
A six-part roadmap is laid out. First come the origins of the memory concept and the token problem, then the classic retrieval approach and graph structures. The second part covers Claude Code's default memory behavior, the third the llm-wiki format, and the fourth tracking with Obsidian. The fifth part asks where ready-made tools like NotebookLM fit into this picture, and the final part examines the built-in memory system of Hermes Agent.
The core problem is one everyone recognizes: after receiving a prompt, the agent reads every file, relevant or not, burns tokens, and keeps no separate record of the user. Each new session re-reads the old files. The video's thesis is that an agent remembering something has exactly one practical meaning: reading the right file at the right time. So the system the user must build is no magic, just a file layout that makes the agent read the correct files on every entry.
The theoretical core is two paths and three parts. The write path synthesizes the finished conversation into files; the read path finds the right file when a question arrives and carries its content back into the chat. The three parts are the data store, the map, and the rule file: a store of Markdown files, an index map telling the agent which file to visit, and the rule file on the Claude Code side defining how to read the map. The video explains it with a city-map analogy: stores are shops, the index is the directions, and the rule file is the guide to reading the map.
The token pain comes from the reading side. If writing is done poorly, the agent reads everything and the bill grows. The read path already exists; what needs building is writing the files so that the map comes out right. A memory system is therefore nothing more than data recorded properly and read back correctly by the agent. This definition becomes the yardstick for every tool choice in the rest of the video.
Next comes the origin of the concept: the retrieval approach. In the classic setup, large datasets were split into chunks, converted into vectors by embedding models, and the nearest chunk was attached to the answer when a question arrived. That is exactly the method the author once used for bulky 3D coordinate files at a previous job. But that setup served static, giant datasets, not a personal system updated every day. In the current agent layout the model itself does the retrieving, and the human's job is to build the index and the map correctly.
The graph concept becomes concrete through the Obsidian interface. Nodes stand for units of work and edges for transfers of information and results; an edge counts only when a genuine transfer happens. The query mechanism grows out of this: the user asks, the agent reads the relevant files, produces a joint result, and that synthesis is saved as a new document and updated in later sessions. The agent is effectively trained from scratch by being shown how to operate the system.
The most candid section covers Claude Code's default behavior. The agent does not truly remember; a subsystem called auto memory keeps notes on expertise, in-chat corrections, ongoing projects, and external references. An index file named MEMORY.md loads partially at the start of each session, while chat records sit in project folders for thirty days by default and are then deleted. The result is familiar: the agent knows who you are but cannot follow a topic from months ago, forgets the format of routine work, and the token problem returns on large projects.
The llm-wiki format is introduced as the solution layer. It is only a Markdown layout standard: vault structure, index map, log records, and cross-links between pages follow a fixed schema. The author builds a live example from topics collected out of viewer comments, handing the agent ingest, query, and health-check commands manually while the agent synthesizes files and updates the index and the log. The striking point is that Obsidian is nowhere in this stage; the format and the commands do the linking work.
Obsidian sits in this architecture as a monitoring and visualization layer, not storage. Opening the vault reveals in graph view which page links where and which nodes are orphaned. Rule files without links show up as orphans, while well-linked pages become routes the agent can follow. The warning is blunt: dumping files into the visualizer before building links with an agent shows nothing but orphan nodes. For one-off small projects, setting up a separate vault is judged unnecessary.
The fifth part takes on the ready-made tool question through NotebookLM. Starting from a fresh account, adding sources, uploading a design file, and generating summaries and info cards are demonstrated. The verdict is balanced: a ready tool is enough for one-off research, but the knowledge lives only inside that app and reaching it costs extra tokens. If the vault is built well, the agent need not read dozens of files; the ready tool and the personal vault are not rivals but answers at different scales.
The final part covers Hermes Agent, and the difference is put in one sentence: Claude Code writes down topics, Hermes writes down operations. A continuously running service, a task triggered every morning, and a skill file acting as the brain check each round whether the conversation holds anything worth learning, then record short notes without touching the main chat. The folder layout and skill mechanism resemble Claude Code, but the write-read-cleanup loop running automatically out of the box sets it apart from the llm-wiki setup. In closing, the author notes the commands she used are shared in the description.
AI commentary
"I see this video as a lesson everyone who works with agents daily should notebook; for the first time someone explains memory not as a magical plugin but as a file system built on write-and-read discipline, this plainly."
AI assessment
Let me steelman the opposing view: perhaps none of this setup is needed. Context windows grow every year, Claude Code's built-in memory already carries session summaries, and dumping everything into context is still the cheapest fix on small projects. That objection holds at small scale, yet independent writing repeats the same warning: using the window as storage degrades performance long before the limit, and old decisions quietly evaporate. A bigger window postpones the problem instead of solving it.
I see two limits in the video's methodology. First, the narrative is almost entirely Claude Code-centric; there is no comparative test against alternative shells or agent frameworks, and the token savings are claimed from experience rather than measured. Second, figures like the thirty-day record retention and the partial loading of the index file are version-sensitive; an independent guide confirms the partial loading behavior, but nothing guarantees it will survive future releases. Unmeasured advice and version-bound detail should be kept apart.
On verifiability the picture is positive. The video's critical figures intersect with independent sources: the partial loading of the index file and auto memory living in project folders are described identically in an independent guide, while NotebookLM source and plan caps are tabulated plan by plan in a current review. My one note on interests: the commands used in the video sit behind links in the description, and such personal templates usually carry traffic toward the author's ecosystem; that does not invalidate the narrative, but viewers should adapt the commands to their own vault rather than adopt them blindly.
My practical verdict is this: for someone working in Claude Code daily, with large projects and repeating routines, this video is worth a setup guide; starting with the llm-wiki layout and monitoring it through Obsidian is a sensible first step. For one-off research work that much infrastructure is overkill, and a ready source-chat tool does the job more cheaply. The sentence I wrote in my own notebook is this: memory is a file-discipline problem, not a model problem.
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @Selma Kocabiyik — video Selma Kocabiyik — video
- @vectorize.io https://vectorize.io/articles/claude-code-memory
- @ianlpaterson.com https://ianlpaterson.com/blog/claude-code-memory-architecture
- @atlasworkspace.ai https://www.atlasworkspace.ai/blog/notebooklm-limitations
- @mem0.ai https://mem0.ai/blog/context-window-is-ram-not-storage-why-most-agent-failures-happen-how-to-fix-them-in-2026
- @pablooliva.de https://pablooliva.de/the-closing-window/obsidian-and-markdown-in-the-ai-agent-era
agent memory · claude code · obsidian · notebooklm · hermes agent · tokens · rag