9 quotes from 1 episode on Latent Space, each with a timestamped link to the source.
9 quotes1 episode
The short version
Alex Zhang states that deploying large amounts of compute on brute-force exploration requires strict verification and expert guidance to prevent capital waste. In unguided agent swarms, 95% of the compute simply explores dead ends and burns tokens.
Most interesting insights
The technology sector now has the capacity to aim $40 million in compute at a single complex problem to achieve a solution.
“It is very exciting that we even have the option to point $40 million at a problem and solve it.”
Frontier laboratories historically conditioned developers to default to standard large language models for every application.
“For the longest time because the labs are the only places that control, you are never going to use something other than GPT or comparable models because they are the best models…”
Exploring dead ends burns tokens without contributing to a final answer. Alex Zhang notes that 95% of the agents in an unguided swarm operate uselessly because researchers take for granted how swarms converge.
“95% of the swarm is entirely useless or like what it's exploring is entirely you're just burning tokens.”
AI-generated GPU kernels suffer from reward hacking that causes the code to fail in production. When evaluated in end-to-end systems, only one kernel out of the top 10 fastest entries actually remained stable.
“GPU kernels have a verification problem. Like we've kind of known this. It's been a problem since kernel bench was released. Like there's a lot of reward hacking that goes on.”
A single engineer with domain knowledge can uncover specific solutions that eliminate the cost of burning a trillion tokens on unstructured model exploration.
“…maybe you can burn like a hundred billion or a trillion tokens on something but if you bring in someone who knows something about the problem um they can uncover something for the model that would like erase that one trillion token spent.”
OpenAI ran an experiment deploying 10,000 agents over 88 hours, consuming 130 billion output tokens with an estimated public API cost of $40 million.
MIT researcher Alex Zhang estimates that 95% of compute in unguided agent swarms is pure waste, exploring dead ends that never contribute to the final answer.
Frontier labs trained the entire industry to assume a language model must always be a token-by-token autoregressive transformer decoder.
Systems like GEV prove that modifying the model output space produces much faster and cheaper inference for targeted tasks like verification, gaming, and classification.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode, and we use it only when a separate check of the captions finds that person on the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.