

π€π¨ Sundae Robotics 10: Robots That See and Act with Their Whole Body
βπ€π¨ Grab a sundae and join Sundae Robotics, a private, invite-only Sunday series bringing together robotics researchers, founders, and builders working at the frontier of physical intelligence.
βSundae Robotics 10
Whole-Body Manipulation, Active Perception & Hardware-Software Co-Design
Featured Talk: Robots That See and Act with Their Whole Body
βRobots That See and Act with Their Whole Body
βKeynote: Xiaomeng Xu (εΎιθ)
Fourth-year PhD Candidate, Stanford University Β· REALab
βFor decades, robot manipulation has largely centered on arms with six or seven degrees of freedom, an end-effector for interacting with objects, and one or two cameras. What becomes possible when we expand these design choices?
βMore joints offer greater flexibility, distributed sensors provide richer perception, soft and growing bodies enable access to confined spaces, and mobility extends the robot workspace. Yet these capabilities also introduce challenges that have traditionally constrained robot design: complex control, unpredictable dynamics, and increasingly large observation and action spaces.
βIn this talk, Xiaomeng will explore how robot learning allows us to embrace this complexity and enable new capabilities.
βShe will introduce RoboPanoptes, a robot that achieves whole-body dexterity through whole-body vision; PanoVine, a system for controlling a soft, growing vine robot through whole-body visual feedback; and HoMMI, which learns whole-body mobile manipulation and active perception directly from robot-free human demonstrations.
βTogether, these systems show how pairing novel hardware capabilities with appropriate learning methods and interfaces enables intuitive teaching, robust control, and scalable data collection for robots that see and act with their whole bodies.
βXiaomeng Xu is a fourth-year PhD candidate in Electrical Engineering at Stanford University, advised by Prof. Shuran Song and part of the REALab. She is supported by the Stanford Interdisciplinary Graduate Fellowship.
βXiaomeng's research focuses on leveraging machine learning algorithms to design and control robots, enabling them to perform complex manipulation tasks efficiently and robustly in the real world. Specifically, she is interested in whole-body manipulation and hardware-software co-design.
βHer recent work explores how unconventional robot embodiments, richer sensing, and learning-based control can enable capabilities that are difficult to achieve with traditional robot designs. RoboPanoptes combines a 9-DoF modular body with 21 distributed cameras to enable whole-body vision and whole-body dexterity. PanoVine uses 19 cameras distributed along a soft, growing vine robot to learn end-to-end whole-body visuomotor control.
βHer work on HoMMI develops a data collection and policy learning framework that learns whole-body mobile manipulation and active perception directly from robot-free human demonstrations, without requiring teleoperation data. HoMMI was published at RSS 2026 and received the Best Paper Award at the Bimanual Robot Learning Workshop at IROS 2026.
βXiaomeng has also worked on active perception, contact-rich manipulation, robot morphology design, bimanual mobile manipulation, and learning from human demonstrations and corrections. Her broader research asks how robot morphology, sensing, learning algorithms, and human interfaces can be designed together rather than treated as independent components.
βBefore Stanford, Xiaomeng received a B.E. in Automation Engineering from Tsinghua University, where she worked with Prof. Li Yi and Prof. Leonidas J. Guibas.
βShe is currently also a research intern with Amazon's Frontier AI & Robotics team, working on humanoid loco-manipulation.
βTopics
ββ’ Robots that see and act with their whole bodies
ββ’ Whole-body manipulation and whole-body dexterity
ββ’ Hardware-software co-design for robot learning
ββ’ Expanding beyond traditional fixed-base robot arms
ββ’ Learning control for high-dimensional and unconventional robot embodiments
ββ’ Distributed sensing and whole-body visual perception
ββ’ RoboPanoptes and whole-body vision with 21 distributed cameras
ββ’ Learning dexterous control for high-degree-of-freedom robots
ββ’ PanoVine and visuomotor control for soft, growing vine robots
ββ’ Using distributed visual feedback to control complex and unpredictable dynamics
ββ’ HoMMI and whole-body mobile manipulation from human demonstrations
ββ’ Robot-free data collection for scalable imitation learning
ββ’ Learning active perception directly from human behavior
ββ’ Mobile manipulation and extending the robot workspace
ββ’ Soft robotics and manipulation in confined spaces
ββ’ Intuitive interfaces for teaching complex robot behaviors
ββ’ Robust control under large observation and action spaces
ββ’ Combining novel hardware capabilities with modern learning methods
βOpen Discussion + Q&A
ββ’ What becomes possible when we stop designing manipulation systems around a single arm and a small number of cameras?
ββ’ How should robot learning methods change as robots gain more joints, sensors, and modes of interaction?
ββ’ When does additional embodiment complexity create useful capability rather than unnecessary control difficulty?
ββ’ Should perception sensors be concentrated in a robot's head or distributed throughout its body?
ββ’ How can a robot learn to decide not only what to do, but also where and how to look?
ββ’ What representations are most effective for learning control from large numbers of distributed cameras?
ββ’ How can learning methods handle the unpredictable dynamics of soft and growing robots?
ββ’ Can unconventional robot morphologies unlock tasks that standard rigid manipulators cannot solve?
ββ’ How should robot morphology, sensing, and policy learning be co-designed rather than optimized independently?
ββ’ Can human demonstrations provide supervision for whole-body behaviors without requiring expensive robot teleoperation?
ββ’ How much of whole-body mobile manipulation can be learned directly from egocentric human data?
ββ’ What information needs to be preserved when transferring behavior from a human embodiment to a robot embodiment?
ββ’ Can active perception emerge naturally from demonstrations, or does it require an explicit objective?
ββ’ How do we scale data collection when observation and action spaces grow dramatically with robot complexity?
ββ’ What does a truly whole-body foundation policy for robotics need to perceive, represent, and control?