

90/30 ML Reading #38: Execution-Grounded AI Research
Execution-Grounded Automated AI Research: From Ideas to Experiments
Join us at Mox to discuss Towards Execution-Grounded Automated AI Research, a new paper exploring how large language models can move beyond generating plausible ideas to actually discover effective algorithms through automated execution.
This work introduces a high-throughput automated idea executor that turns natural-language research ideas into runnable code, launches large-scale GPU experiments, and feeds execution results back to models as learning signals. Using realistic environments for LLM pre-training and post-training, the authors show that frontier models can successfully implement and evaluate a large fraction of their own ideas, enabling closed-loop, execution-grounded research at scale.
We’ll dive into how execution-guided evolutionary search rapidly discovers methods that outperform strong baselines (e.g., beating GRPO on post-training and halving nanoGPT pre-training time), why this approach is far more sample-efficient than best-of-N prompting, and how models begin to generate genuinely algorithmic, not just hyperparameter, ideas. We’ll also examine a key negative result: why reinforcement learning from execution reward improves average idea quality but collapses diversity and fails to push the frontier.
🔎 Analyzed Paper
Towards Execution-Grounded Automated AI Research
Discussion at 20:00, (optional) quiet reading from 19:00.