arXiv:2601.10553cs.CV2026-01被引 23

用潜在世界模型提升视频生成的物理合理性,推理时即可优化。

Inference-time Physics Alignment of Video Generative Models with Latent World Models

  • 用潜伏世界模型作奖励,在推理阶段优化生成轨迹。
  • 在多条件生成中显著提升物理合理性,人类偏好测试验证有效。
  • 在ICCV 2025 PhysicsIQ挑战赛中以62.64%得分夺冠,领先7.42%。

当前最先进的视频生成模型虽能产出视觉上吸引人的内容,但常违背基本物理规律,限制其实际应用。尽管部分研究归因于预训练阶段缺乏物理理解,我们发现推理策略不当也是导致物理合理性不足的关键原因。为此,本文提出WMReward,将提升视频生成的物理合理性视为一个推理时对齐问题。具体而言,利用潜伏世界模型(此处为VJEPA-2)强大的物理先验作为奖励信号,搜索并引导多个候选去噪轨迹,从而实现测试时计算资源的扩展,提升生成质量。实验表明,该方法在图像条件、多帧条件和文本条件生成场景中均显著改善了物理合理性,且经由人类偏好研究验证。尤为突出的是,在ICCV 2025 Perception Test PhysicsIQ挑战赛中,取得62.64%的最终得分,排名第一,较此前最优结果高出7.42%。本工作证明了使用潜伏世界模型提升视频生成物理合理性的可行性,不局限于特定实例或参数化形式。

原文摘要 · Abstract (English)

State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency to insufficient physics understanding from pre-training, we find that the shortfall in physics plausibility also stems from suboptimal inference strategies. We therefore introduce WMReward and treat improving physics plausibility of video generation as an inference-time alignment problem. In particular, we leverage the strong physics prior of a latent world model (here, VJEPA-2) as a reward to search and steer multiple candidate denoising trajectories, enabling scaling test-time compute for better generation performance. Empirically, our approach substantially improves physics plausibility across image-conditioned, multiframe-conditioned, and text-conditioned generation settings, with validation from human preference study. Notably, in the ICCV 2025 Perception Test PhysicsIQ Challenge, we achieve a final score of 62.64%, winning first place and outperforming the previous state of the art by 7.42%. Our work demonstrates the viability of using latent world models to improve physics plausibility of video generation, beyond this specific instantiation or parameterization.

视频生成物理对齐潜伏世界模型推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。