Video Models Are Zero-Shot Learners and Reasoners by Thaddäus Wiedemer
Abstract: The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. Could video models be on a trajectory towards general-purpose vision understanding, much like LLMs developed general-purpose language understanding? We demonstrated that Veo 3 can zero-shot solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, completing symmetric patterns, solving mazes, and much more. These emergent abilities indicate that video models are on a path to becoming unified, generalist vision foundation models.
Speaker: Thaddäus Wiedemer — Research Scientist at Google, whose work on zero-shot reasoning in video models helped define the field of video reasoning.
⚠️ Special time: 10:00 AM CEST · 4:00 PM Beijing · 1:00 AM PT (moved to daytime for our Europe + China audience)
Website: https://journal.video-reason.com/
Register on this page to receive the Zoom join link (sent automatically with your confirmation).
Subscribe to our mailing list for weekly talk announcements: Open Google Form