

Of Foundations and Superstructures: Why World Modeling Loves On-Task Data - Roy Fox
Zoom Link:
https://nyu.zoom.us/j/97237939451?pwd=08Yl50yARiKAhc7NCJubRb4MtAa77q.1
Meeting ID: 972 3793 9451
Passcode: 640434
Abstract: An agent learning to control its environment would often benefit from modeling it. Compared to control policies, world models use a richer training signal and tend to generalize and transfer better. However, for a world model to induce good behavior, it must be highly accurate in all reachable states, which may require too much data. Because efforts to leverage web-scale data for control are yet to succeed as they famously have for vision and language, we ask: can information encoded in vision and language foundation models help guide world modeling? In the first part of this talk, we will see two such methods: one that uses a segmentation foundation model to block visual distractions and keep state representations task-relevant; and one that queries a language model to hypothesize about abstract world models that guide exploration and planning. In the second part of the talk, we will revisit the transfer power of world models in two settings: simulation-to-reality and delayed perception. We will see how a model of a simulator can be adapted to reality with a tiny amount of data; and how a world model can transfer across varying delays of the agent's observations. Throughout, we will ponder how these successes inform the challenges of control foundation models and ways they could be overcome.
Bio: Roy Fox is an Assistant Professor of Computer Science at the University of California, Irvine. His research interests include theory and applications of control learning: reinforcement learning (RL), control theory, information theory, and robotics. His current research focuses on structured and model-based RL, language for RL and RL for language, robot safety, and optimization in deep control learning of virtual and physical agents.