让视频生成懂物理:通过隐式学习物理规律提升运动真实性
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

- 联合建模视觉内容与隐式物理动态,不依赖显式物理公式
- 在标准与物理感知基准上均超越现有方法,保持视觉质量
- 适合需要真实物理运动的视频生成场景,如仿真、动画
近年来,生成视频模型在大规模数据和强大架构推动下取得了显著的视觉真实感。然而,单纯扩大数据和模型规模并未赋予系统对真实世界动态底层物理规律的理解。现有方法常无法捕捉或强制实现物理一致性,导致运动不自然。本文提出Phantom,一种融合物理信息的视频生成模型,通过联合建模视觉内容与隐式物理动态,利用物理感知的视频表示作为底层物理的抽象嵌入,在无需显式指定复杂物理动态与属性的情况下,联合预测物理动态并生成未来帧。该设计使生成视频兼具视觉真实性和物理一致性。在标准视频生成与物理感知基准上的定量与定性结果表明,Phantom不仅在物理一致性上优于现有方法,且在感知保真度方面表现相当。
原文摘要 · Abstract (English)
Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However, emerging evidence suggests that simply scaling data and model size does not endow these systems with an understanding of the underlying physical laws that govern real-world dynamics. Existing approaches often fail to capture or enforce such physical consistency, resulting in unrealistic motion and dynamics. In his work, we investigate whether integrating the inference of latent physical properties directly into the video generation process can equip models with the ability to produce physically plausible videos. To this end, we propose Phantom, a Physics-Infused Video Generation model that jointly models the visual content and latent physical dynamics. Conditioned on observed video frames and inferred physical states, Phantom jointly predicts latent physical dynamics and generates future video frames. Phantom leverages a physics-aware video representation that serves as an abstract yet informaive embedding of the underlying physics, facilitating the joint prediction of physical dynamics alongside video content without requiring an explicit specification of a complex set of physical dynamics and properties. By integrating the inference of physical-aware video representation directly into the video generation process, Phantom produces video sequences that are both visually realistic and physically consistent. Quantitative and qualitative results on both standard video generation and physics-aware benchmarks demonstrate that Phantom not only outperforms existing methods in terms of adherence to physical dynamics but also delivers competitive perceptual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。