πŸ€–πŸ¨ Sundae Robotics 05: 3D World Models, Spatial Intelligence & Adaptive AI β€” Modality Forcing & 3D Pre-training Objectives

Hosted by Edmond & 4 others
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

β€‹πŸ€–πŸ¨ Grab a sundae and join Sundae Robotics, a private, invite-only Sunday series bringing together robotics researchers, founders, and builders working at the frontier of physical intelligence.

​Sundae Robotics 05
3D World Models, Spatial Intelligence & Adaptive AI
Featured Talk: Modality Forcing & 3D Pre-training Objectives

​Understanding the 3D World: Scalable Pre- and Post-training Recipes for Adaptive AI

​Keynote: Bardienus (Bart) Duisterhof
Final-year PhD Student, Carnegie Mellon University Robotics Institute (Jeffrey Ichnowski) Β· World Labs Β· Collaborator with Deva Ramanan

​How do we build generative models that deeply understand the 3D world, and adapt quickly to new tasks? In this talk, Bart will present his recent contributions in scalable pre-training and post-training recipes. First, he will present a new pre-training objective for learning 3D dynamics. Inspired by masked representation learning, the approach explores scalable supervision for learning representations that transfer to downstream world modeling and imitation learning.

​Second, Bart will cover Modality Forcing, a post-training recipe to extract spatial knowledge from image generation models. Modality Forcing trains a single DiT to model the joint distribution between modalities by setting a separate noise level for each modality. The method competes with the very best depth models and scales with text-to-image pre-training.

​Finally, Bart will discuss ongoing work toward more scalable and adaptive AI. On the pre-training side, he will argue that video generation is wasteful, motivating the search for more compressed transition models. For post-training, he will lay out directions toward adaptation across tasks at a fraction of today's cost. Together, these directions aim at a future where anyone can adapt powerful open models to new tasks.

​Bart's research develops generative models for spatial intelligence, spanning 3D/4D representation learning and scalable pre- and post-training recipes that transfer to downstream tasks. He is a final-year PhD student at Carnegie Mellon University's Robotics Institute, where he is advised by Jeffrey Ichnowski and collaborates with Deva Ramanan. He currently works on 3D pre-training at World Labs. His recent work includes Modality Forcing, which showed that text-to-image pre-training substantially improves 3D perception. Bart is increasingly interested in the foundations of generation β€” how these models learn, and how to make them more adaptable and efficient. He believes in a future with powerful open models that can be quickly adapted for new tasks, and is passionate about making open AI beneficial for society.

​Pre-Reading

​‒ Modality Forcing: Extracting spatial knowledge from image generation models
https://modality-forcing.github.io

​Topics

​‒ Scalable pre-training for 3D world understanding

​‒ Learning transferable representations of 3D dynamics

​‒ Extracting spatial knowledge from image generation models

​‒ 3D/4D representation learning for world modeling and imitation learning

​‒ Scaling spatial intelligence with text-to-image pre-training

​‒ Compressed transition models beyond video generation

​‒ Efficient post-training and adaptation across downstream tasks

​‒ Open and adaptable foundation models

​Open Discussion + Q&A

​‒ What is the right pre-training objective for models that need to understand and act in the 3D world?

​‒ Can point tracks provide a more scalable representation for world modeling than pixels or video?

​‒ How much spatial intelligence is already latent inside large image generation models?

​‒ Is video generation an unnecessarily expensive way to learn world dynamics?

​‒ How can powerful foundation models be adapted to new tasks at a fraction of today's cost?

Location
84 Serrano Dr
Atherton, CA 94027, USA