Key Takeaways
- Local AI agents store context in local markdown files, but ephemeral cloud containers destroy that file system state on restart.
- Nick Kuhn from VMware Tanzu Platform argues that agent architectures must follow 12-factor app principles by cleanly separating compute from state.
- Decoupling state into a dedicated memory service lets systems scale to 1,000 container instances without losing conversational history or operational progress.
- Team-scoped memory services allow ephemeral agents to spin up on demand, inherit historical architecture and roadmap data, and shut down cleanly.
The Local File System Trap
Most developers start building AI agents on their laptops. The setup is simple: the agent writes notes, project state, and task lists into local markdown files on the hard drive. On a local machine, that persistent scratchpad works without issues.
Moving that setup into production breaks immediately. Production cloud systems run on ephemeral containers that boot, handle work, and terminate. Nick Kuhn pointed out the failure mode during his conversation on Practical AI: “The traditional harnesses or agents were kind of built originally just to be like, 'Oh, I've got file system access, and I can just write a bunch of MD files, and that's like my memory, everything's great.' You're like, 'Well, you start to get into cloud cloud world that things spin up and down and are ephemeral, you're going to want to save those MD files somewhere. They're going... They're going to go go off into the ether.'”
When a cloud container crashes or scales down to zero, any context written to the container's local disk disappears. If your agent relies on local file writes for its memory, your production rollout loses its state the moment traffic shifts or a node recycles.
12-Factor Rules for Agent Architecture
The solution to ephemeral container state is not new; platform engineers solved it over a decade ago with 12-factor application design. The rule is straightforward: treat compute as disposable and store all state in backing services.
Kuhn connects this cloud heritage directly to agent systems: “where like you talk about Cloud Foundry and Tanzu Platform, it was kind of based on this premise of a 12-factor application, where like storage and state are a little, you know, decoupled cleanly from the actual apps, you could like, you know, I can I can scale up to a thousand instances and be fine, and you handle session state and things.”
If you decouple state from your agent runtime, the container becomes purely an execution engine. It accepts a prompt, pulls external context, runs tool calls, updates the external memory store, and exits. You can run one instance or scale up to a thousand parallel workers without session corruption or missing files.
Scoping Memory Across Team Agents
Decoupled memory also changes how agents collaborate across an organization. Instead of locking context inside isolated single-user runtimes, you attach a shared memory service scoped by team and domain.
Under this pattern, an agent spinning up for a quick background task does not start from a blank prompt. “So as they come up, they know like all of this historical data of like what like the architecture of the application, the, you know, what what's been done before and all like the roadmap, that type of thing,” Kuhn noted. As local agents do work, the shared memory service grows, allowing any fresh instance to inherit that institutional context instantly.
What to Do With This
Audit your agent codebase tomorrow morning and search for local file writes, sqlite files, or raw markdown scratchpads on the local disk. Replace those disk writes with an API client that pushes structured context and project state to an external backing store. Verify that you can destroy your agent container mid-task, boot a brand new container on a different host, and resume execution without missing context.