通过融合视觉与物理感知,生成更符合真实物理规律的视频。
MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

- 将语义、几何和时空轨迹统一为伪RGB格式,让模型理解物理动态。
- 在Videophy和PhyGenbench上显著提升物理合理性,优于现有方法。
- 适合需要真实物理行为的视频生成研究者或应用开发者。
尽管视频扩散模型能生成视觉效果出色的视频,但常因仅基于像素重建而产生物理不一致的结果。为此,我们提出MMPhysVideo,首次通过联合多模态建模提升视频生成的物理合理性。我们将语义、几何及时空轨迹等感知线索转化为统一的伪RGB格式,使视频扩散模型能够直接捕捉复杂物理动态。为减少跨模态干扰,设计双向控制教师架构,采用并行分支解耦RGB与感知处理,并通过两个零初始化控制连接逐步建立像素级一致性。为提高推理效率,利用表征对齐将教师模型的物理先验蒸馏至单流学生模型。此外,提出MMPhysPipe端到端数据构建与标注管道,借助视觉-语言模型结合视觉证据链规则定位物理主体,驱动专家模型提取多粒度感知信息。无需额外推理开销,MMPhysVideo在Videophy和PhyGenbench基准上持续提升先进模型的物理合理性,性能超越现有方法。
原文摘要 · Abstract (English)
Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first study to enhance physical plausibility in video generation through joint multimodal modeling. We recast perceptual cues, specifically semantics, geometry, and spatio-temporal trajectories, into a unified pseudo-RGB format, enabling VDMs to directly capture complex physical dynamics. To mitigate cross-modal interference, we propose a Bidirectionally Controlled Teacher architecture, which utilizes parallel branches to fully decouple RGB and perception processing and adopts two zero-initialized control links to gradually establish pixel-wise consistency. For inference efficiency, the teacher's physical prior is distilled into a single-stream student model via representation alignment. Furthermore, we present MMPhysPipe, an end-to-end data curation and annotation pipeline tailored for constructing physics-rich multimodal datasets. MMPhysPipe employs a vision-language model (VLM) guided by a chain-of-visual-evidence rule to pinpoint physical subjects, enabling expert models to extract multi-granular perceptual information. Without additional inference costs, MMPhysVideo consistently improves physical plausibility of advanced models on the Videophy and PhyGenbench benchmarks and achieves superior performance among existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。