Key Takeaways
- Reinforcement learning workloads now match or exceed pre-training compute volume, forcing data centers to shuttle jobs between different specialized chips instead of relying on uniform GPU clusters.
- Specialized inference chips like Groq LPUs, Cerebras wafer-scale engines, and AWS Trainium are entering production pipelines alongside Nvidia hardware.
- Huawei produces over 1 million AI chips per year across its Ascend 910B, 910C, 940, 950, and 960 series, while competing Chinese chipmakers remain capped in the hundreds of thousands.
- Frontier Chinese labs are already running production workloads on domestic hardware, with GLM deploying 50,000 non-Nvidia Chinese accelerators for inference on its newest model.
The End of the Monolithic GPU Cluster
Building AI infrastructure used to mean buying identical Nvidia GPUs, wiring them together, and running one uniform cluster. That setup is cracking under modern post-training workloads. Reinforcement learning requires massive inference passes to generate tokens, test answers, and update policy networks. As swyx pointed out during the conversation, “RL is like as big or sometimes bigger than the pre-trained workload right now, right? And so that actually means that you need like maybe a new form of disagregation where like you have the different kinds of compute in the same data center that you shuttle back.”
Training requires heavy compute density and high memory bandwidth for backward passes. Inference requires low latency and high batch throughput per dollar. Running both stages on the same expensive Nvidia flagship GPUs burns cash. Data centers are splitting the pipeline across multi-silicon environments. Teams now pair Nvidia training clusters with specialized inference hardware like Groq LPUs, Cerebras systems, or AWS Trainium. SemiAnalysis researcher Jordan noted that this heterogeneous shift is arriving immediately: “Multisilicon going to be huge next year like with the verbin generation with Nvidia we're going to see it because they're going to bring the LPUs in from croc the LPX systems right so even if we're just testing Nvidia we're going to have multisilicon to test.”
Huawei and the Chinese Silicon Supply Race
While western clouds mix domestic hardware options, Chinese labs face strict export controls that block access to western supply chains. That pressure forced rapid domestic development. Following initial delays after US trade restrictions, Huawei restructured its semiconductor operations and moved from its Ascend 910B and 910C series into rapid releases of the Ascend 940, 950, and 960.
SemiAnalysis researcher Dylan Patel outlined the production reality in China: “If you look at the Chinese companies, really, Huawei is the only one that even in gross chips, they're making a million plus. And then the Huawei chips are behind. So then you have to discount them for that. Everyone else is still in the hundreds of thousands of units.”
Chinese AI companies are not waiting for performance parity. They are deploying domestic hardware directly into live production. As Patel noted, “GLM said explicitly right that they had 50,000 chips that were not Nvidia. They were Chinese-made AI accelerators for inference of their newest model.” When an engineering team cannot buy western hardware, they optimize software runtimes, rewrite kernels, and split model architectures across thousands of lower-yield chips.
What to Do With This
Audit your model deployment pipeline this week. If you run reinforcement learning rollouts or high-volume inference on the same Nvidia H100 or B200 instances you use for training, separate them. Benchmark your token-generation stage on AWS Trainium or dedicated inference ASICs like Groq, and calculate your cost per million tokens across both setups before signing your next compute contract.