Key Takeaways

  • OpenAI partnered with Broadcom to design its first in-house ASIC, codenamed Jalapeño, engineered strictly for large language model inference.
  • The custom chip delivers 1.5x to 1.9x greater inference throughput per watt than Nvidia GB200 and GB300 systems.
  • Jalapeño cuts end-to-end latency between 1.7x and 3.6x compared to standard GPUs.
  • OpenAI aims to deploy 10 gigawatts of these custom accelerator systems between late 2026 and 2029, multiplying its existing compute footprint by five.

The Shift From Training to Pure Inference

First-generation silicon usually disappoints. Building custom chips requires immense capital, deep semiconductor design talent, and years of iteration before matching merchant silicon. Yet OpenAI took a different path by focusing entirely on inference rather than training.

General-purpose graphics cards like Nvidia H100s or B200s carry architectural overhead to support arbitrary training workloads, matrix multiplications, and dynamic developer requirements. An application-specific integrated circuit (ASIC) discards that baggage. As John Coogan observed on TBPN when reviewing Dylan Patel's analysis, “Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Reuben.”

By stripping away features needed only for pre-training, OpenAI and Broadcom tuned Jalapeño around memory bandwidth and token generation. Coogan highlighted the performance gains: “Roughly 1.5 to 1.9x more useful inference throughput per watt than Nvidia GB200, GB300 systems while simultaneously cutting end-to-end latency 1.7 to 3.6x.” For interactive products like ChatGPT and voice interfaces, raw token latency dictates product feel just as much as model intelligence.

Why Efficiency Dictates AI Gross Margins

Every foundation model lab faces an energy wall. Data center capacity cannot expand fast enough to meet user demand if every query burns through hundreds of watts of power on general-purpose chips. Building dedicated silicon allows OpenAI to serve more users within tight physical grid limits.

Coogan framed the business reality clearly: “Good to stretch that energy further. You need less data centers to do more inference. That's good. Good for gross margins.” When energy availability caps your operational ceiling, throughput per watt directly determines how many paying customers you can support.

OpenAI plans to roll out these chips at a scale rarely seen in private technology infrastructure. Coogan detailed the roadmap: “10 gigawatts of open AI designed accelerator systems manufactured by Broadcom or in partnership with Broadcom deployed from the second half of 2026.”

Jordi Hays put the massive scale into perspective: “Put it into context. That's roughly five times more compute than the whole company currently has operational roughly.” Shifting that volume to proprietary chips reduces OpenAI's reliance on Nvidia's pricing power while locking in structural unit-economic advantages for the rest of the decade.

What to Do With This

Audit your model infrastructure bills today. Calculate your exact split between training spend and inference spend over the last two quarters. If inference accounts for more than 70 percent of your compute costs, test specialized inference endpoints like Groq, Cerebras, or provider-native ASICs this week to see if you can cut token latency and double your gross margins without changing your model weights.