arXiv:2603.03485cs.CVcs.AI2026-03

让视频生成模型学会真实物理规律,生成更逼真的4D动态场景。

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

  • 分三阶段训练:先预训练几何与运动,再用仿真数据微调,最后用强化学习修正物理错误。
  • 在长时序下保持几何一致和运动稳定,相比基线提升显著的物理合理性。
  • 适合做高保真4D建模、虚拟世界生成的研究者与开发者使用。

近期的视频扩散模型虽在大规模生成方面表现卓越,但常缺乏精细的物理一致性,导致时间演化中出现不合理的动态。本文提出Phys4D,一种从视频扩散模型学习物理一致4D世界表征的流程。该方法采用三阶段训练范式:首先通过大规模伪监督预训练建立稳健的几何与运动基础;其次利用仿真生成数据进行基于物理的监督微调,确保4D动态的时间一致性;最后应用仿真引导的强化学习,修正难以通过显式监督捕捉的残余物理违规。为评估细粒度物理一致性,我们引入一套4D世界一致性评估,涵盖几何一致性、运动稳定性及长时序物理合理性。实验表明,相较于仅注重外观的基线模型,Phys4D在时空与物理一致性上均有显著提升,同时保持强大生成能力。

原文摘要 · Abstract (English)

Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present \textbf{Phys4D}, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts \textbf{a three-stage training paradigm} that progressively lifts appearance-driven video diffusion models into physics-consistent 4D world representations. We first bootstrap robust geometry and motion representations through large-scale pseudo-supervised pretraining, establishing a foundation for 4D scene modeling. We then perform physics-grounded supervised fine-tuning using simulation-generated data, enforcing temporally consistent 4D dynamics. Finally, we apply simulation-grounded reinforcement learning to correct residual physical violations that are difficult to capture through explicit supervision. To evaluate fine-grained physical consistency beyond appearance-based metrics, we introduce a set of \textbf{4D world consistency evaluation} that probe geometric coherence, motion stability, and long-horizon physical plausibility. Experimental results demonstrate that Phys4D substantially improves fine-grained spatiotemporal and physical consistency compared to appearance-driven baselines, while maintaining strong generative performance. Our project page is available at https://sensational-brioche-7657e7.netlify.app/

4D建模物理一致性视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。