Cover Image for πŸ€–πŸ¨ Sundae Robotics 07: Long-Horizon Robot Manipulation, Hierarchical Control & World Models β€” Retriever
Cover Image for πŸ€–πŸ¨ Sundae Robotics 07: Long-Horizon Robot Manipulation, Hierarchical Control & World Models β€” Retriever
38 Going

πŸ€–πŸ¨ Sundae Robotics 07: Long-Horizon Robot Manipulation, Hierarchical Control & World Models β€” Retriever

Hosted by Edmond, Mene Mazarakis & James (Jingxi) Xu
Register to See Address
Atherton, CA
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

β€‹πŸ€–πŸ¨ Grab a sundae and join Sundae Robotics, a private, invite-only Sunday series bringing together robotics researchers, founders, and builders working at the frontier of physical intelligence.

​Sundae Robotics 07
Long-Horizon Robot Manipulation, Hierarchical Control & World Models
Featured Talk: Retriever β€” Closing the Perception-Reasoning-Action Loop for Long-Horizon Robot Manipulation

​Retriever: Closing the Perception-Reasoning-Action Loop for Long-Horizon Robot Manipulation

​Keynote: Linfeng Zhao
Postdoctoral Scholar, Stanford University Β· Stanford Robotics Center

​Robots performing long-horizon tasks must connect perception, reasoning, and action while coordinating asynchronous processes and maintaining state across modules and time scales. In this talk, Linfeng will present Retriever, a programming framework developed for asynchronous, hierarchical robot systems.

​Retriever coordinates slower components, such as vision-language planning and memory, with faster perception and VLA policy/control loops without forcing every module onto a single clock. Using object search and multi-step manipulation as examples, Linfeng will show how this design supports closed-loop execution.

​He will also discuss his recent work on manipulation policy learning and world modeling.

​Linfeng Zhao is a Postdoctoral Scholar at Stanford University, working with Prof. Mykel Kochenderfer and Prof. Jeannette Bohg at the Stanford Robotics Center.

​He finished his Ph.D. in 2025 at the Khoury College of Computer Sciences at Northeastern University, advised by Prof. Lawson L.S. Wong, where he closely collaborated with the MIT Learning and Intelligent Systems group, Prof. Leslie Kaelbling, and Prof. Robin Walters.

​He has interned at Meta, the Boston Dynamics AI Institute, Amazon, and Microsoft Research Asia. Before that, he worked with Prof. Hao Su at UC San Diego from 2018–2019.

​Linfeng's research focuses on building human-level general-purpose agents that can act in the physical worldβ€”robots that navigate homes, manipulate objects, and accomplish long-horizon tasks in open-world scenarios with unseen environments and goals.

​He develops abstractions for decision-making that decompose complex behaviors into compositional building blocks. He also develops learning and planning approaches to enable agents to reason about the world and plan their actions at decision time for scalable, generalizable, and efficient decision-making systems.

​Topics

​‒ Closing the perception-reasoning-action loop for long-horizon robot manipulation

​‒ Asynchronous, hierarchical robot systems

​‒ Coordinating vision-language planning, memory, perception, and control

​‒ Integrating slow reasoning modules with fast VLA policy and control loops

​‒ Maintaining state across modules and time scales

​‒ Programming abstractions for complex robot behavior

​‒ Object search in open-world environments

​‒ Multi-step manipulation and closed-loop execution

​‒ Compositional abstractions for decision-making

​‒ Decision-time reasoning and planning

​‒ Manipulation policy learning

​‒ World modeling for robot decision-making

​‒ General-purpose agents for unseen environments and goals

​‒ Scalable, generalizable, and efficient robot decision-making systems

​Open Discussion + Q&A

​‒ How should robot systems coordinate components that naturally operate at very different time scales?

​‒ Should perception, planning, memory, and control share a unified architecture, or remain modular?

​‒ What state needs to persist across a long-horizon manipulation task?

​‒ How should high-level vision-language reasoning interact with low-level VLA policies?

​‒ When should a robot re-plan versus continue executing its current policy?

​‒ What abstractions make long-horizon robot behaviors easier to compose and debug?

​‒ How can asynchronous architectures avoid stale observations, plans, or memories?

​‒ What does closed-loop execution require beyond simply calling a planner repeatedly?

​‒ How should robots recover when an intermediate step in a multi-stage task fails?

​‒ Can hierarchical systems generalize more effectively to unseen environments and goals?

​‒ What role should world models play in decision-time robot planning?

​‒ How should manipulation policies and explicit reasoning systems divide responsibility?

​‒ Can modular robot systems scale to genuinely open-world household tasks?

​‒ What would a practical architecture for a human-level general-purpose physical agent look like?

Location
Please register to see the exact location of this event.
Atherton, CA
38 Going