Issue No. 40Week ending Sunday, October 4, 2026522 episodes · 2294 articles
The Throughline ↓
The Podcast Summary.

10+ hours of podcasts, in 5 minutes.

Which GPU Clouds Are Actually Good? | ClusterMAX 3.0

With swyx, Jordan · Sunday, October 4, 2026

SemiAnalysis researchers Jordan and Dylan join swyx to break down the findings of ClusterMAX 3.0, evaluating the managed GPU cloud landscape. They analyze severe cybersecurity vulnerabilities in NeoClouds, extreme compute shortages driven by profitable inference providers, Nvidia's defensive acquisition strategy with Poolside and Hugging Face, and the future of multi-silicon and Chinese AI hardware.

Key takeaways

  • Google possesses massive TPU capacity, yet internal teams like DeepMind have less dedicated R&D compute than OpenAI and Anthropic. Read more →
  • Frontier labs like Anthropic and OpenAI used to rent GPU clusters in blocks of 8,000 chips; now they aggressively secure slices as small as 1,000 GPUs. Read more →
  • Reinforcement learning workloads now match or exceed pre-training compute volume, forcing data centers to shuttle jobs between different specialized chips instead of relying on uniform GPU clusters. Read more →
  • SemiAnalysis built ClusterMAX 3.0 to test managed software layers rather than raw data center construction or bare-metal server specs. Read more →
  • Nvidia pulls in nearly $50 billion of free cash flow each quarter, giving its corporate strategy team roughly $200 billion annually to deploy defensively. Read more →
  • SemiAnalysis audits of GPU NeoClouds revealed that researchers could view other tenants' private training runs, storage volumes, and active workloads. Read more →

6 articles from this episode

More Latent Space episodes

Every Latent Space episode we cover →

The Sunday Email

Get next Sunday's issue in your inbox.

10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday.

Newsletters

For now, every subscriber gets both newsletters. No spam. Unsubscribe with one click.