Key Takeaways
- Illuminate builds reinforcement learning environments inside Docker containers to train frontier AI models on non-coding knowledge work.
- Training setups give agents raw data, Excel workbooks, and PowerPoint decks, scoring them on completed tasks like building an LBO or a sensitivity table.
- Jerry Woo observes a "Moore's law of RL environments," where the complexity of training gyms doubles every 6 to 8 months.
- The long-term goal of RL gyms extends beyond single spreadsheets to full simulations of companies, industries, and governments.
The Docker Classroom for White-Collar Work
Software engineers have compilers, unit tests, and terminal outputs. When an AI writes code, the machine gives an immediate, deterministic answer: it runs, or it breaks. White-collar knowledge work has never had that clean feedback loop. A financial analyst building a merger model or an associate structuring a slide deck relies on fuzzy human review. That makes reinforcement learning difficult.
Jerry Woo and his team at Illuminate approach this by building synthetic playgrounds for non-coding tasks. Instead of grading static text, they put AI agents into isolated software containers and make them do actual desk work.
“What we do at Illuminate is we build the training benchmarks and training environments used by Frontier Labs to improve their models,” Woo explains. “Specifically we specialize in improving models on non-coding knowledge work.”
Woo views his team as environment builders setting up interactive challenges for models:
Inside that container, the model interacts with the software, executes steps, tests outputs, and receives reward signals. Some tasks are scored through deterministic validation rules, while others use learned reward models. The model learns by trial and error in a sandbox before ever seeing production data.
The Rapid Doubling of Training Complexity
Training environments cannot stay static because frontier models saturate simple benchmarks quickly. A gym that challenges a model today becomes obsolete once the model masters the underlying workflows.
Woo tracks this shift internally through a metric his team uses to plan environment roadmaps. “We internally at Illuminate have a phrase we call the Moore's law of RL environments,” Woo says. “Every six to eight months we've noticed that the complexity of the environment or gym used to train a frontier model roughly doubles.”
Right now, complexity means handling dirty datasets, multi-tab spreadsheets, and slide creation across multiple steps. The next phase will demand coordinated work across entire organizational systems.
“In the future we're going to be building simulations of whole companies, whole governments, whole industries to teach models how to operate autonomously within our societies, within our companies,” Woo explains. Moving from single-agent tasks to multi-agent economic simulations is where training data will have to go as frontier labs push toward autonomous operations.
What to Do With This
Audit the most repetitive 10-step analytical workflow on your team this week, such as updating a weekly financial dashboard or running customer churn cohorts. Package the starting CSV files, the desired output sheet, and three explicit error-checking validation formulas into a clean folder. If you cannot specify deterministic success criteria for your own human workflow, no AI agent will be able to complete it reliably either.