Key Takeaways

  • Asking an LLM to police its own boundaries fails; compliance rules belong in deterministic, non-AI code outside the model context window.
  • Gavriel Cohen built NanoClaw to isolate every agent inside a hardened Docker container sandbox on a single virtual machine within the customer's private cloud perimeter (AWS Bedrock, GCP Vertex, or Azure).
  • An external agent gateway intercepts outbound traffic to enforce routing policies and inject API credentials, blocking raw secret exposure to the model.
  • Approval gates work invisibly: agents send normal payloads without knowing a review queue exists, waiting while the orchestrator surfaces proposed actions to human operators.

The Method

Most enterprise AI demos fail security reviews because founders build safety rules into system prompts. Cohen took the opposite route with NanoClaw after turning a 40-hour weekend project into an enterprise platform. If you tell a model to ask for permission before running a bash command or accessing customer data, prompt injection or drift will eventually bypass that request.

Cohen moved all enforcement out of the model entirely: “Agents run in sandboxes. All requests are proxied through the agent gateway and are then, that's where policies are enforced and credentials are injected.”

First, isolate execution at the operating system layer. In Cohen's architecture, one virtual machine runs multiple agents, but each agent sits inside its own hardened Docker container. As Cohen explains, “one Nanoclaw deployment, one instance of Nanoclaw running in a VM, can run many different agents, and each of those agents are segregated in and isolated in their own sandbox, which is a hardened Docker container, and they don't have any access to information that other agents have access to.”

Second, make approvals invisible to the agent. When an agent drafts an outbound message or an action, it does not prompt the user for permission. It simply transmits the message as standard output. “That approval is not It's not a prompt for the agent saying you need to ask for approval. The agent doesn't even know that there's an approval process. It just sends messages back and forth the way it normally does,” Cohen notes.

Third, wire strict agent-to-agent boundaries into the deterministic orchestration layer. Non-AI code inspects message targets before routing them. “With those wirings of which agent can talk to which agent and where there are these approval gates, you can enforce data access policies and control and prevent an agent that has access to really sensitive data from sharing that information with an agent that has access to the internet,” Cohen says. If an agent with internal database access wants to pass records to an outward-facing research agent, the orchestrator halts the payload and pings a human reviewer.

Where This Breaks Down

This architecture adds latency to multi-agent loops. When you introduce external gateway proxies and human approval gates, tasks that once ran in seconds can stall for hours waiting on human review. If you create too many gates, human operators develop permission fatigue, clicking approve without reading the payload. The model also breaks down if your engineers bypass the gateway during local testing by mounting host volumes or baking production API keys directly into local Docker environment files.

What to Do With This

Audit your agent pipeline tomorrow morning. Strip every sentence from your system prompts that instructs the model to request user permission before taking an action. Replace those prompts with a deterministic proxy function in your API router that pauses tool execution and posts the raw JSON payload to a dedicated Slack approval channel before executing the call.