Key Takeaways

  • OpenAI introduced 12-hour prompt cache guarantees to slash costs for persistent agent loops that outlive standard 30-minute eviction windows.
  • The pre-warming API lets developers pay prompt write fees in advance, keeping large context blocks hot before user traffic arrives.
  • Developers can automate context reduction through token-threshold compaction in the Responses API or trigger it manually with /compact.
  • Nikunj Handa compares OpenAI's current platform trajectory to Stripe, building higher-level infrastructure on top of raw inference calls.

The Economics of Long-Running Context

Agents that run for hours break standard API economics. Every round trip in a continuous session resends tool schemas, conversation history, and system instructions. If that context falls out of memory between actions, you pay full write and compute costs on every single turn.

OpenAI previously capped cache retention at short 30-minute intervals. That worked for instant chat sessions, but multi-hour workflows suffered constant cache misses.

“We provide now guarantees of cache hits within 30 minutes,” Handa explains. “For one of our users, we just launched a much longer cache window. We have a 12-hour caching guarantee that we offer so that you have guaranteed cache hits.”

For builders running asynchronous workflows, this turns previously prohibitive token bills into predictable fixed costs.

Pre-Warming and Context Compaction

Predictable caching requires two things: keeping context alive when it is idle, and pruning it before it exhausts the window.

To solve the cold-start problem, OpenAI added pre-warming directly to the API. Handa outlines the mechanic: “If you know that you are going to get this prompt, you can pre-warm the cache, pay the cache write fee right now, and then have it ready to go for the next 30 minutes.”

Once a thread runs long enough, context bloat eventually degrades model performance. OpenAI offers two distinct paths to handle pruning in its developer stack:

“If you're in the Responses API, there are two ways of doing it,” Handa says. “One is what we call server-side compaction, which is you basically tell the Responses API that if you ever hit this threshold of tokens, just auto-compact it and reduce the context being used. And the second way is /compact, which is if you want full control so you can compact at any time and have your own logic.”

Server-side compaction acts as an automatic safety valve for standard agent loops. Manual compaction gives teams with complex memory requirements the ability to selectively summarize or drop specific tool outputs without losing critical system state.

Primitives Above the Base Layer

Handling cache eviction and context pruning directly inside the API represents a shift in how model providers package infrastructure. Handa previously worked at Stripe, and he sees direct parallels in how AI platform architectures are maturing.

“At Stripe, a lot of the game was building these higher-level primitives and products on top of the core payments primitives,” Handa notes. “I'm always curious about what the best way of doing that is in AI.”

Instead of forcing developers to build custom Redis state trackers and summarization middleware, model platforms are baking state persistence into the protocol.

What to Do With This

Review your active agent pipelines this week. Identify loops that pause for more than 30 minutes between actions, and configure pre-warming for your largest static prompt prefixes. If your agent uses custom truncation scripts, set a target token threshold in the Responses API to offload compaction directly to the server.