

What Happens When AI Runs the Robotics Experiments?
Jia Qi Yip (Lead Research Engineer, Menlo Research) will share on “Agentic Autoresearch for Reinforcement Learning Locomotion.”
Training reinforcement learning policies in simulation still involves a lot of manual trial and error. Researchers tune rewards, wait for training runs, benchmark the resulting policies, and decide what to try next. But what happens when an AI agent takes on that research loop itself?
Jia Qi will share what happened when they handed this outer loop to an agent for humanoid locomotion on Asimov, an open-source humanoid robot. Working unattended for 16 to 32 hours at a time, the agent used Isaac Lab for training and a separate MuJoCo benchmark as its judge. It ran reward ablations, architecture sweeps, and recurrent (LSTM) policies, and even built its own terrain benchmark with stairs, slopes, and obstacles.
Its best policy reduced path drift by more than half compared with their production baseline, without losing push robustness. Jia Qi will unpack what the agent got right, where it still needed a human, and what this means for how reinforcement learning research is done.
More About the Speaker
Jia Qi Yip is the lead research engineer at Menlo Research, where he works across the full stack of Asimov, an open-source humanoid robot, from speech and agents to locomotion, manipulation and hardware. His focus is closing the robotics deployment gap by applying the latest research to tough real-world problems.
More About the Series
Singapore Embodied AI Colloquium (SEAIC) is a focused technical research series built around researcher-driven discussions and sharing on frontier topics in embodied AI, robot learning, and physical AI.
Get involved: Learn more about Lorong AI | Speaker Sign-up | WhatsApp Community | LinkedIn | X