Key Takeaways
- SemiAnalysis audits of GPU NeoClouds revealed that researchers could view other tenants' private training runs, storage volumes, and active workloads.
- In one audit, researchers could see confidential intelligence workloads run by military and national security agencies on shared infrastructure.
- The security failures do not require frontier model attacks; they stem from two- to three-year-old unpatched CVEs and misconfigured InfiniBand partition keys.
- Public warnings from AI research leaders like Ilya Sutskever match these findings: GPU clouds often ignore standard infrastructure isolation.
- SemiAnalysis released a free CLI tool called CMAX to help engineering teams check cluster software versions and flag outdated drivers.
Leaking National Intelligence on Shared Hardware
When startups spin up thousands of GPUs on alternative cloud providers to dodge waitlists and high prices, they assume basic multi-tenant isolation works. It often does not.
SemiAnalysis audited multiple GPU providers during their ClusterMAX evaluations. What they discovered went far beyond minor software bugs. Dylan Patel noted that basic isolation controls were missing across several providers: “In one case, we were literally able to see national intelligence of a certain country's stuff. They had multiple national security agencies and military style intelligence stuff running in that NeoCloud, and it's like good god, we're just here to test the cluster and we can see this stuff.”
If independent researchers testing cluster throughput can inspect national intelligence workloads, an adversary with dedicated compute access can capture your proprietary model weights, fine-tuning datasets, and internal prompts without breaking modern encryption.
Basic Keys and Three-Year-Old CVEs
Silicon Valley spends billions securing application layers while the physical network plumbing in AI clouds sits wide open. The breaches observed in these audits did not involve zero-day exploits or elite nation-state toolkits. They were ordinary operational failures.
“The checks are very simple,” Jordan explained. “Is your software version up to date? Are your drivers up to date? Is it configured correctly? These are relatively simple checks, but when people are running versions of software that are two to three years old and they have well-documented CVEs online, you don't need a frontier model to create a exploit of this.”
The vulnerability extends directly into high-speed networking fabric. High-performance GPU clusters rely on InfiniBand and specialized Ethernet configurations for distributed training. InfiniBand uses management keys (M-keys) and partition keys (P-keys) to enforce network boundaries between different tenants. When operators fail to configure these keys, tenant boundaries dissolve. Patel pointed out: “InfiniBand and Ethernet have various controls. At least on the InfiniBand side, it's M-keys, P-keys, all these things. Are those properly configured and deployed? It's shocking how bad most NeoClouds are.”
AI leaders like Ilya Sutskever have raised alarms about physical compute security for frontier AI labs. These audit findings show those warnings reflect current reality rather than hypothetical risks. When cloud providers race to rack GPUs and bill hours, baseline security operations get skipped.
What to Do With This
Run the free CMAX CLI tool against your active GPU instances tomorrow morning to audit installed drivers, kernel patches, and known CVEs. If your cloud provider cannot verify isolated InfiniBand P-keys and strict tenant network segmentation in writing, assume every training job, dataset, and model checkpoint on that cluster is accessible to adjacent tenants.