Cover Image for Do Joint Audio-Video Generation Models Understand Physics? by Zijun Cui
Cover Image for Do Joint Audio-Video Generation Models Understand Physics? by Zijun Cui
Avatar for Video Model Journal Club
Hosted By
21 Went

Do Joint Audio-Video Generation Models Understand Physics? by Zijun Cui

Zoom
Registration
Past Event
Welcome! To join the event, please register below.
About Event

​Abstract: Recent joint audio-video generation models such as Seedance 2.0, Kling 3.0 Omni, and Veo 3.1 can produce clips that look and sound very realistic. In the real world, however, vision and sound are two observations of the same physical event. Turning up a volume knob makes the music louder, and a ringing alarm clock sealed in a foam-lined box sounds muffled. Do current models understand the physics that connects what we see with what we hear, or do they simply put together plausible frames and sounds?

​This talk introduces AV-Phys Bench, the first comprehensive benchmark for physical commonsense in joint audio-video generation. The benchmark contains prompts organized by how a scene evolves — Steady State, Event Transition, and Environment Transition — together with an Anti-AV-Physics set where prompts deliberately ask for physical violations. The prompts instantiate 41 audio-visual physics principles, and each prompt comes with its own rubric that scores semantic adherence and physical commonsense within and across modalities. Evaluating 3 proprietary and 4 open-source models with about 58,000 human judgments reveals a consistent gap between semantics and physics, much weaker performance on transition scenes, and a sharp drop of 45% to 69% on anti-physics prompts.

​The talk also presents AV-Phys Agent, a ReAct-style evaluator that grounds a multimodal LLM with deterministic audio DSP measurements such as loudness, pitch, reverberation, and onset timing. It aligns with human ratings more closely than MLLM-as-judge baselines and enables scalable physics-aware evaluation, and closes with thoughts on using such evaluators as verifiable reward signals for post-training.

​Speaker: Zijun Cui is a Ph.D. student in Computer Science at the University of Texas at Dallas, advised by Prof. Yapeng Tian in the Computer Vision and Multimodal Computing (CVMC) Lab. Zijun works on Physical AI with a focus on audio-visual perception and generation, and is the lead author of AV-Phys Bench, a benchmark for physical commonsense in joint audio-video generation. Zijun is also a co-first author of SAVVY (NeurIPS 2025 oral), which studies spatial reasoning with audio-visual LLMs. Before joining UT Dallas, Zijun received an M.S. in Electrical and Computer Engineering from the University of Washington and a B.Eng. from Zhejiang University.

​All sessions: https://luma.com/video-model Register on this page to receive the Zoom join link (sent automatically with your confirmation).

​Subscribe to our mailing list for weekly talk announcements: Open Google Form

Avatar for Video Model Journal Club
Hosted By
21 Went