

π€π¨ Sundae Robotics 07: Long-Horizon Robot Manipulation, Hierarchical Control & World Models β Retriever
βπ€π¨ Grab a sundae and join Sundae Robotics, a private, invite-only Sunday series bringing together robotics researchers, founders, and builders working at the frontier of physical intelligence.
βSundae Robotics 07
Long-Horizon Robot Manipulation, Hierarchical Control & World Models
Featured Talk: Retriever β Closing the Perception-Reasoning-Action Loop for Long-Horizon Robot Manipulation
βRetriever: Closing the Perception-Reasoning-Action Loop for Long-Horizon Robot Manipulation
βKeynote: Linfeng Zhao
Postdoctoral Scholar, Stanford University Β· Stanford Robotics Center
βRobots performing long-horizon tasks must connect perception, reasoning, and action while coordinating asynchronous processes and maintaining state across modules and time scales. In this talk, Linfeng will present Retriever, a programming framework developed for asynchronous, hierarchical robot systems.
βRetriever coordinates slower components, such as vision-language planning and memory, with faster perception and VLA policy/control loops without forcing every module onto a single clock. Using object search and multi-step manipulation as examples, Linfeng will show how this design supports closed-loop execution.
βHe will also discuss his recent work on manipulation policy learning and world modeling.
βLinfeng Zhao is a Postdoctoral Scholar at Stanford University, working with Prof. Mykel Kochenderfer and Prof. Jeannette Bohg at the Stanford Robotics Center.
βHe finished his Ph.D. in 2025 at the Khoury College of Computer Sciences at Northeastern University, advised by Prof. Lawson L.S. Wong, where he closely collaborated with the MIT Learning and Intelligent Systems group, Prof. Leslie Kaelbling, and Prof. Robin Walters.
βHe has interned at Meta, the Boston Dynamics AI Institute, Amazon, and Microsoft Research Asia. Before that, he worked with Prof. Hao Su at UC San Diego from 2018β2019.
βLinfeng's research focuses on building human-level general-purpose agents that can act in the physical worldβrobots that navigate homes, manipulate objects, and accomplish long-horizon tasks in open-world scenarios with unseen environments and goals.
βHe develops abstractions for decision-making that decompose complex behaviors into compositional building blocks. He also develops learning and planning approaches to enable agents to reason about the world and plan their actions at decision time for scalable, generalizable, and efficient decision-making systems.
βTopics
ββ’ Closing the perception-reasoning-action loop for long-horizon robot manipulation
ββ’ Asynchronous, hierarchical robot systems
ββ’ Coordinating vision-language planning, memory, perception, and control
ββ’ Integrating slow reasoning modules with fast VLA policy and control loops
ββ’ Maintaining state across modules and time scales
ββ’ Programming abstractions for complex robot behavior
ββ’ Object search in open-world environments
ββ’ Multi-step manipulation and closed-loop execution
ββ’ Compositional abstractions for decision-making
ββ’ Decision-time reasoning and planning
ββ’ Manipulation policy learning
ββ’ World modeling for robot decision-making
ββ’ General-purpose agents for unseen environments and goals
ββ’ Scalable, generalizable, and efficient robot decision-making systems
βOpen Discussion + Q&A
ββ’ How should robot systems coordinate components that naturally operate at very different time scales?
ββ’ Should perception, planning, memory, and control share a unified architecture, or remain modular?
ββ’ What state needs to persist across a long-horizon manipulation task?
ββ’ How should high-level vision-language reasoning interact with low-level VLA policies?
ββ’ When should a robot re-plan versus continue executing its current policy?
ββ’ What abstractions make long-horizon robot behaviors easier to compose and debug?
ββ’ How can asynchronous architectures avoid stale observations, plans, or memories?
ββ’ What does closed-loop execution require beyond simply calling a planner repeatedly?
ββ’ How should robots recover when an intermediate step in a multi-stage task fails?
ββ’ Can hierarchical systems generalize more effectively to unseen environments and goals?
ββ’ What role should world models play in decision-time robot planning?
ββ’ How should manipulation policies and explicit reasoning systems divide responsibility?
ββ’ Can modular robot systems scale to genuinely open-world household tasks?
ββ’ What would a practical architecture for a human-level general-purpose physical agent look like?