Key Takeaways

  • Synchronous tool calls create artificial latency bottlenecks; async function calling lets the model reason and generate continuously while external operations execute.
  • OpenAI moved its real-time API layer to bidirectional WebSockets, replacing the traditional request-response loop with an open channel for continuous streaming.
  • Mid-turn steering allows developers to inject instructions and new context directly into an active reasoning trace before the model finishes its turn.
  • Optimization loops for UltraFast inference are automated internally, using autonomous Codex agents targeting Astra to squeeze out latency.

Stop Blocking the Reasoning Loop

Building agents on top of classic HTTP endpoints forces a rigid loop: prompt the model, halt execution, run a function, append the result, and prompt the model again. If an external database query or browser action takes four seconds, your application sits idle for four seconds.

Nikunj Handa explained that OpenAI rebuilt this interaction pattern for newer systems: “What you see with like a lot of the things that you're seeing in like Codex and Dots and everything is that tool calls take so long that you don't have to pause the model's execution while the tool is running. So you could just kick off a tool call, keep running, keep reasoning, and then check back in.”

Decoupling execution from inference treats tool calls as background jobs rather than blocking gates. When an agent requires three slow API calls, it can dispatch all three, continue outlining its plan or processing other subtasks, and ingest the payloads whenever the network returns them.

Bidirectional WebSockets Enable Live Steering

Async execution only works if you can communicate with the model mid-stream. Under the standard REST architecture, a model turn is atomic: once triggered, you wait for completion. OpenAI shifted to WebSockets to make model inference fully interactive.

“WebSockets just opens this whole bidirectional communication thing with the model,” Handa noted. This protocol change enables mid-turn steering, giving developers a direct hook into the active context window. “We launched mid-turn steering. So now you can inject messages while the model is reasoning in the middle. As your tool call finishes, you can put in that instructions.”

If a background tool fails or returns unexpected data, you no longer need to cancel the run, pay for wasted output tokens, and restart. You inject a correction into the open stream while the model is still thinking. The model absorbs the new direction immediately and adjusts its current trajectory.

Agents Optimizing Inference Stacks

Reducing protocol latency exposes raw model inference as the next bottleneck. OpenAI paired this architectural update with UltraFast inference, built and refined using automated engineering loops.

Handa described how the infrastructure team tunes throughput: “The most fun part of UltraFast has been just watching the inference team cook with Astra. They are constantly having these Codex agents running trying to squeeze out more performance.”

Instead of human engineers manually benchmarking every kernel tweak, autonomous Codex agents write, test, and profile optimizations against Astra directly. This machine-driven loop shrinks the round-trip latency budget, making continuous steering responsive enough for real-time production workloads.

What to Do With This

Audit your primary agent pipeline tomorrow and flag every synchronous tool call that takes longer than 500 milliseconds. Refactor the slowest integration into an asynchronous background task over WebSockets, letting your agent continue planning while the task resolves. Test mid-turn steering on your longest workflow to inject tool results directly into active generation instead of restarting the prompt loop.