Cover Image for Claas Voelcker - On-policy value learning at 10000 frames per second
Cover Image for Claas Voelcker - On-policy value learning at 10000 frames per second
Led by Rahul Narava and Gusti Winata. Part of the Cohere Labs Open Science initiative https://cohere.com/research/open-science
Hosted By

Claas Voelcker - On-policy value learning at 10000 frames per second

Google Meet
Registration
Past Event
Welcome! To join the event, please register below.
About Event

Talk Description: When samples are cheap and fast to collect, the RL community relies on the policy-gradient theorem to obtain agents which can reliably train on massive amounts of data. However, since their inception, zeroth-order algorithms such as REINFORCE, TRPO, and PPO have been plagued by high variance, which makes them hard to tune. In the off-policy regime where sample efficiency is the goal, stable and efficient value-driven methods have been explored, but these require replay buffers and specialized architectures to stabilize off-policy learning.

What if we could bridge the two paradigms, and bring stable value-driven learning to the on-policy sampling regime? In our new ICLR paper, Relative Entropy Policy Optimization, we explore how to achieve stable on-policy value function learning at 50000 frames per second. We will see that value-learning is possible, and useful, without the use of massive replay buffers by combining insights from both the on-policy and off-policy literature.

Bio: Claas is a PostDoc working at the intersection of Reinforcement and robotics at the University of Texas at Austin with Peter Stone and Amy Zhang. He holds a PhD from the University of Toronto and the Vector Institute, where he was advised by Profs. Amir-massoud Farahmand and Igor Gilitschenski. Outside of research, Claas is also a core organizer at Queer in AI, an affinity group that builds community for queer researchers and industry practitioners.

His research focuses on stabilizing value estimation and representation learning for efficient and reliable reinforcement learning in robotics. He is driven by the question of how we can learn to accurately predict the value of taking actions, a central task in RL. Beyond that, he investigates how we can do better science in RL by thinking about what problems we should be benchmarking our exciting advances on.

Led by Rahul Narava and Gusti Winata. Part of the Cohere Labs Open Science initiative https://cohere.com/research/open-science
Hosted By