让视频生成模型更符合物理规律,提升仿真可靠性。
PhyWorld: Physics-Faithful World Model for Video Generation

- 分两阶段微调:先优化连续帧一致性,再用物理偏好引导
- 在物理忠实度评测中得分3.09,优于基线2.99
- 适合用于训练物理感知的AI系统,安全且可扩展
世界模拟器可在真实部署前为物理人工智能系统提供安全、可扩展的训练环境。大型视频生成模型因其能生成多样且逼真的视觉未来,正成为此类模拟器的有前途基础。然而,将其作为世界模拟器需生成与输入条件一致且符合基本物理规律的视频延续。本文提出PhyWorld,一种通过两阶段后训练实现时间连贯且物理忠实的场景延续的视频生成世界模型。第一阶段通过光流匹配微调提升视频到视频延续的稳定性与运动一致性;第二阶段利用直接偏好优化(DPO)对物理偏好样本进行对齐,引导模型生成更具物理合理性的输出。我们采用标准视频质量基准和专用物理忠实度基准(含逐定律评分)评估。实验显示,PhyWorld在VBench上平均得分为0.769,优于最先进基线(≤0.756);在物理忠实度基准上平均得分为3.09,高于最强基线2.99。结果表明,通过延续信号与物理偏好信号后训练的大规模视频生成模型,可显著提升其作为物理人工智能世界模拟器的有效性。
原文摘要 · Abstract (English)
World simulators can provide safe and scalable environments for training Physical AI systems before real-world deployment. Large video generation models are emerging as a promising basis for such simulators because they can generate diverse and realistic visual futures. However, using them as world simulators requires physically faithful video continuations, namely, generated videos that preserve the physical state implied by the conditioning input, and evolve in ways consistent with basic physical principles. We propose PhyWorld, a video generation world model designed to produce temporally coherent and physically faithful scene continuations through two-stage post-training. In the first stage, we improve video-to-video continuation with flow matching fine-tuning, encouraging stable visual attributes and coherent motion dynamics across frames. In the second stage, we align generated dynamics with physical principles using Direct Preference Optimization (DPO) over physics preference pairs, guiding the model toward outputs with higher physical plausibility. To evaluate PhyWorld, we use both standard video-quality benchmarks and a dedicated physical-faithfulness benchmark with per-law scoring. Experiments show that PhyWorld improves video consistency, achieving an average score of 0.769 on VBench compared with 0.756 or below for state-of-the-art baselines. PhyWorld also improves physical plausibility, reaching an average score of 3.09 on our physical-faithfulness benchmark compared with 2.99 for the strongest baseline. These results suggest that post-training large video generation models with continuation and physics-preference signals can make them more effective world simulators for Physical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。