Key Takeaways

  • OpenAI built the Decisions API in under four weeks without training a single new model weight.
  • The system runs directly on existing Luna weights with reasoning paths disabled to maximize raw throughput.
  • The architecture combines constrained structured outputs, aggressive time to first token (TTFT) tuning, and parallel question evaluation.
  • The primary production targets are high-speed support ticket triage and real-time voice tool calling with GPT Live.

The Method

Most teams assume building a specialized, ultra-fast classification endpoint requires fine-tuning or training a custom model from scratch. OpenAI took the opposite path. Nikunj Handa explained that his team delivered the Decisions API in less than a month by repurposing what they already had.

“So we haven't trained like a new model for this. We're like building this purely on top of the same Luna weights that we have,” Handa said.

Instead of launching a training run, an engineer from inference and an engineer from infrastructure teamed up to strip away overhead. They turned off reasoning steps and rebuilt the serving layer around three clear tactical mechanics.

First, they locked the output space down with rigid schemas. Handa noted: “On top of that, what you're doing is you're constraining. So like structured output is a big part of it. You're really optimizing the inference stack to get very fast on TTFT.” Constraining the token vocabulary reduces decode time and guarantees strict schema adherence on every request.

Second, they introduced parallel question execution over shared contexts. When a system needs to evaluate five distinct boolean conditions or categorization tags on a single input document, standard API setups make five sequential calls or run one large, slow prompt. The Decisions API runs those discrete evaluations in parallel over the same prompt cache, slicing total latency.

Third, they tuned the engine for real-time agent loops. Classification speed matters most when an agent sits in an interactive execution loop. Handa pointed to voice-driven workflows as the immediate beneficiary: “People have been putting together these tool calling demos of GPT Live controlling a computer and just feels like so much more snappy and natural. So I'm kind of excited to see what people do with Live and with Luna on Decisions API.”

Where This Breaks Down

Stripping reasoning to hill-climb on raw latency works only for narrow, deterministic judgments.

If your classification task requires multi-step deductive logic, math, or resolving contradictory context, this architecture will produce hallucinations. The model cannot backtrack or deliberate. It generates immediate, constrained probabilities based on pattern matching across the context window.

Similarly, if your application requires rich qualitative explanations alongside classifications, the constrained schema approach falls apart. Adding open-ended text fields ruins the speed gains achieved by tiny token budgets.

What to Do With This

Audit your primary user-facing agent loop this week. Identify every classification or routing call currently running on standard reasoning-heavy models like GPT-4o. Replace those sequential routing prompts with constrained structured output schemas that run parallel evaluations across fixed enum values. If you are waiting on multi-second API calls just to route a user request or select a tool, you are wasting user patience on compute you do not need.