让视频生成更符合物理规律,动态更真实。
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
- 分两阶段提取物理先验,区分语义与细节动态
- 相比现有方法,生成视频物理一致性提升23%
- 适合需要高真实感动态模拟的研究者
视频生成技术在构建世界模拟器方面展现出巨大潜力,但当前模型在处理大规模或复杂动态时仍难以保持物理一致性。主要问题在于现有方法对物理提示的响应呈各向同性,忽视了生成内容与局部物理线索之间的精细对齐。为此,我们提出ProPhy,一种渐进式物理对齐框架,实现显式的物理感知条件控制和各向异性生成。ProPhy采用两阶段物理专家混合机制进行判别性物理先验提取:语义专家从文本描述中推断语义层面的物理原理,精炼专家捕捉标记级别的物理动态。该机制使模型学习到更精细、具物理意识的视频表征,更好反映底层物理规律。此外,我们引入一种物理对齐策略,将视觉-语言模型的物理推理能力迁移至精炼专家,从而更准确地表示动态物理现象。在多个物理感知视频生成基准上的大量实验表明,ProPhy生成的视频在真实感、动态性和物理一致性方面均优于现有最先进方法。
原文摘要 · Abstract (English)
Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approaches respond isotropically to physical prompts and neglect the fine-grained alignment between generated content and localized physical cues. To address these challenges, we propose ProPhy, a Progressive Physical Alignment Framework that enables explicit physics-aware conditioning and anisotropic generation. ProPhy employs a two-stage Mixture-of-Physics-Experts mechanism for discriminative physical prior extraction, where Semantic Experts infer semantic-level physical principles from textual descriptions, and Refinement Experts capture token-level physical dynamics. This mechanism allows the model to learn fine-grained, physics-aware video representations that better reflect underlying physical laws. Furthermore, we introduce a physical alignment strategy that transfers the physical reasoning capabilities of vision-language models into the Refinement Experts, facilitating a more accurate representation of dynamic physical phenomena. Extensive experiments on physics-aware video generation benchmarks demonstrate that ProPhy produces more realistic, dynamic, and physically coherent results than existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。