Key Takeaways

  • Anthropic built Claude Mods to run directly inside Claude Code's TypeScript runtime instead of relying on external shell script hooks.
  • Forked sub-agents reuse the parent agent's prompt cache, making targeted supervisory checks like quizzes or assumption reviews virtually free in token cost.
  • Developers can modify both the underlying execution loop and the terminal user interface programmatically during an active session.
  • This architecture marks the arrival of mutable software, where generative applications alter their own tools and interfaces on demand.

In-Process Code Beats External CLI Hooks

Most developer tooling treats extensibility as an afterthought, offering bash hooks that fire on pre-commit or post-build. Anthropic took a different route with Claude Code. They built Claude Mods to execute directly inside the TypeScript client process.

Shihipar explained the internal evolution of the system: “Internally, we were originally calling this function hooks. And so that gives you a little bit of an idea where hooks register an event to happen and then a script to call. And this inside of the TypeScript runtime is running things.”

Running inside the runtime changes what an extension can touch. Instead of just intercepting shell commands, a mod can parse live object state, manipulate the terminal display, and intercept model decisions before they execute. Shihipar noted that “you can customize both the execution and the UI.” The result is software that reshapes its own interface to match the task at hand.

Prompt Caching Makes Supervisory Agents Cheap

The hardest problem in autonomous coding is preventing the model from running wild on bad assumptions. The standard fix is human oversight, but stopping for every minor decision destroys developer flow.

Claude Mods solves this by spinning up lightweight forked sub-agents to act as supervisors. You can run assumption checkers or automated quizzes that test code changes before applying them. Because of Anthropic's prompt caching architecture, branching an agent does not mean repaying the full context window cost.

Shihipar highlighted the economics: “Basically a forked agent maintains the prompt cache. It is one of those unintuitive things where you can fork and do a little request and it will be very cheap because the entire prompt cache is done.”

When branched requests cost a fraction of a cent, you can spend compute lavishly on self-verification. Shihipar pointed out that “the ways you spend compute are to keep the user in the loop and make sure that you are getting to the right decision and the right output and artifacts, and mods are this way of spending that intelligence.”

The Shift to Mutable Software

We are moving away from rigid binaries toward software that rewires itself while you use it. When an agent can write, test, and load its own extensions on the fly, the boundary between the developer and the tool disappears.

Shihipar views this as an early glimpse of what comes next: “I do think that this is a preview of mutable software, and how generative software you can customize safely.”

If an agent encounters a proprietary API or an internal database format, it does not need a human engineer to build an integration. It can write a mod, mount it in-process, and continue working without restarting the environment.

What to Do With This

Audit your current agent workflows this week. Identify the three most common failure modes where your coding agent makes incorrect assumptions about your codebase. Write a lightweight supervisory mod in TypeScript that forks a sub-agent to validate those specific assumptions against your schema before any file edits are committed.