Key Takeaways

  • Computer use ran ten times faster year-over-year by cutting reasoning overhead, parallelizing execution steps, and speeding up safety checks.
  • Outer scaffolding code must be built ahead of current models and then deleted as newer models make hardcoded logic counterproductive.
  • OpenAI runs a dual-agent safety design where a secondary guardian agent called auto review supervises every live click and action of the primary agent.
  • Proactive background agents maintain continuous context by staying online to research and plan, turning human users into live directors rather than manual prompters.
  • ChatGPT Spaces acts as persistent institutional memory, allowing teams to share collective notes and state across multiple autonomous workers.

The Scaffolding Trap: Why You Must Delete Your Code

Engineers building on top of frontier models often make the mistake of treating their outer runtime wrappers as permanent architecture. Sottiaux points out that this relationship is cyclical and fragile.

When a model struggles with structured planning or tool calling, developers write hundreds of lines of deterministic code, prompt chains, and retry logic to patch the gaps. The problem starts when the next model version drops. If you keep the old scaffolding around, your application runs slower and dumber than raw model calls.

Sottiaux explains that builders must always construct their outer wrappers slightly ahead of current model capabilities to compensate for missing features. Once the next generation catches up, teams have to delete parts of that wrapper because legacy logic holds the model back. Writing code destined for the trash bin feels unnatural to traditional software developers, but treating agent runtime code as temporary scaffolding is the only way to capture frontier model gains without adding latency.

The Guardian Pattern and the 10x Speedup

Making models execute tasks on a computer desktop requires solving two conflicting constraints: raw execution speed and absolute safety. If an agent stops to deliberate before every mouse click or keystroke, tasks take hours. If it moves instantly without guardrails, it can delete databases or leak credentials.

OpenAI achieved a 10x year-over-year speedup in computer use by separating these responsibilities into distinct systems. Instead of forcing one model to do both heavy reasoning and continuous safety verification, they run parallel execution loops alongside a dedicated supervisor.

As Sottiaux explains: “So we don't just click around and let the primary agent do things like we have a second agent which we call auto review or internally we call it guardian which watches over this primary agent.”

The guardian model evaluates safety in real time while the primary worker executes actions. This keeps latency down without sacrificing control. Sottiaux notes that these agents do not merely react to prompts; they stay online continuously, occasionally sleeping, while running background research and aligning with user preferences. Spaces provides the shared memory layer across the entire organization, ensuring that autonomous workers retain institutional context across sessions.

What to Do With This

Audit your existing agent codebase this week and flag every hardcoded retry loop, validation script, and multi-step prompt chain you wrote to fix previous model weaknesses. Run an A/B test with your current frontier model using raw tool calling against your custom wrapper. If the raw model matches or beats your custom wrapper on accuracy, delete that scaffolding code immediately to eliminate latency.