Key Takeaways
- SemiAnalysis built ClusterMAX 3.0 to test managed software layers rather than raw data center construction or bare-metal server specs.
- Crusoe secured large contracted gigawatt pipelines and excels at physical builds, yet scored lower on ClusterMAX due to weak managed cluster software.
- The primary operational test for AI training workloads is whether a cloud provider can automatically hot-swap failed GPU nodes inside Slurm or Kubernetes without crashing the job.
- Buying compute requires evaluating software reliability rather than physical expansion; ClusterMAX is built for engineers spending budgets, not retail stock pickers.
Power pipelines do not fix broken software
Building a data center and running an AI cluster are two entirely different engineering disciplines. A provider can pour concrete, lock down utility substations, and contract gigawatts of power faster than anyone else in the market. That proves mastery over real estate and electrical engineering. It does not prove their software can keep eight thousand GPUs running a training job overnight.
SemiAnalysis encountered this exact split when scoring providers for ClusterMAX 3.0. Crusoe stands out as one of the fastest builders in the sector. Dylan Patel noted their physical strength directly: “Crusoe, it was downgraded for managed clusters. But they're one of the best data center builders in the world, if not the best. They have nearly the most contracted or the most contracted gigawatts in their pipeline.”
Yet when evaluated on the software layer that developers actually interact with, the score dropped. Bare-metal availability is only the starting line. Patel explained: “A lot of these providers like Crusoe have fantastic bare metal. But that's not what we're testing in ClusterMAX. We're not testing bare metal. It is part of the criteria, but a lot of the criteria is a managed cluster: managed Slurm, managed Kubernetes, all of that sort of stuff.”
What AI engineers actually pay for
When teams spin up large model training runs, raw hardware is not enough. Hardware breaks constantly at scale. GPUs drop off the PCIe bus, optical transceivers fail, and networking links drop packets. If your cloud vendor gives you bare metal and walks away, your internal infrastructure team spends their entire week babysitting dead nodes and writing recovery scripts.
That failure recovery mechanism is where many Neoclouds stumble. Automatic node replacement and job rescheduling separate a functional managed compute platform from raw rack rentals. If a provider cannot detect a bad node, pull it out of the job, substitute a warm spare, and resume checkpointing without manual human intervention, your effective cost per GPU hour skyrockets.
Research for buyers, not stock speculators
The surge in AI infrastructure has created an industry of marketing where data center announcements get mistaken for software readiness. Financial markets treat contracted gigawatts as revenue certainty. But engineers tasked with training models need to know if the software actually runs today.
Jordan emphasized the strict purpose of their rankings: “This is strictly for buyers. This is for helping people buy compute. It is not for people to pick what stock to yolo all their savings into.”
Patel added that credibility in infrastructure analysis requires separating marketing claims from operational realities: “The value we have to the industry is that we tell the truth and what we believe. If we didn't do that, then no one would care what we would say and we'd be screaming into the void. So our reputation is what matters the most over a quick buck.”
What to Do With This
Before signing an annual GPU reservation, demand a live test of node-failure recovery. Force a simulated node failure during a distributed Slurm run on at least 64 GPUs. Time how long the control plane takes to detect the fault, isolate the node, hot-swap a spare, and restart from your last checkpoint.