Key Takeaways
- OpenAI introduced three execution tiers for its developer models: normal mode, fast mode at 2x speed, and ULTRAFAST mode at 8x speed.
- GPT-6 Astra ULTRAFAST charges a 6x price premium over standard pricing to achieve sub-second execution.
- Claire Vo built "the other pencil," a collaborative sketchpad where the model generates matching SVG drawings in real time alongside human input.
- Running a dynamically generated 3D game called Little Starship with live physics updates cost Vo $97 in API fees for 30 minutes of play.
Eight Times Faster at Six Times the Price
Speed changes what software can do. When AI responses take four seconds, developers build chatbots and asynchronous summarizers. When latency drops to milliseconds, software turns into a live canvas.
Following OpenAI DevDay 2026, product leader Claire Vo tested the developer tiers available for the new GPT-6 Astra model. As Vo explained, “You can run your models in codecs on normal mode, on fast mode which is two times as fast, or on ultra fast which is eight times as fast. Now it cost six times as much. It's expensive.”
To test whether 8x speed actually unlocks new product surfaces, Vo built two working prototypes instead of running synthetic benchmarks. The first was "the other pencil," an interactive human-AI sketchpad. When a user draws wave shapes on the screen, Astra evaluates the strokes live and draws matching SVG vector graphics right behind the cursor. The interaction feels like two people drawing on the same physical pad of paper at the same table.
“If you get very high intelligence at almost real-time experience, you can build some AI things that I think you couldn't build before,” Vo noted. Sub-second response times remove the waiting state that makes agentic interfaces feel clunky.
The Real-Time 3D Experiment and the $194 Hourly Rate
Vo pushed the model further by building an interactive 3D browser game titled Little Starship. In the game, players modify the virtual world and its underlying physics through direct text prompts while flying. Vo described the architecture: “What happens is it uses Astra ULTRAFAST to render these new 3D objects super fast real time here in the app.”
Generating 3D assets, scene graphs, and motion physics on the fly proved technically viable. However, continuous token generation at top speed came with a severe commercial cost.
“Now, real talk, it cost me like $97 or something to run this for 30 minutes. So, it was not cheap. This is a very expensive game,” Vo admitted. That translates to an operating cost of roughly $194 per hour for a single active player session.
For enterprise workflows or high-value design tools, a $3-per-minute execution cost is easy to justify. If an architect or an industrial engineer saves three days of CAD modeling in a fifteen-minute session, the return on investment is obvious. But for consumer software, games, and mass-market products, building continuous real-time model loops on ULTRAFAST mode will blow through operating margins in an afternoon.
What to Do With This
Audit your product roadmap for features you killed solely because API latency exceeded two seconds. Prototype the highest-value concept on a fast execution tier this week, but place a hard token-spend cap of $10 per session in your middleware before exposing it to users.