用自监督模型提升视频生成的物理合理性,效果提升6%。
Improving the Physics of Video Generation with VJEPA-2 Reward Signal
- 用VJEPA-2作为奖励信号指导生成过程。
- 在PhysicsIQ基准上,物理合理性提升约6%。
- 适合关注视频生成真实性与物理一致性研究者。
本文介绍在ICCV 2025感知测试工作坊举办的PhysicsIQ挑战赛中的优胜方案。当前主流视频生成模型在物理理解方面严重不足,常生成不合理的视频。Physics IQ基准已表明视觉真实不等于物理合理。然而,直观物理理解可通过自然视频的自监督学习(SSL)预训练涌现。本文研究是否可利用基于SSL的视频世界模型来提升视频生成模型的物理合理性。具体地,我们基于先进的视频生成模型MAGI-1,结合新提出的视频联合嵌入预测架构2(VJEPA-2),将其作为生成过程的奖励信号。结果表明,通过使用VJEPA-2作为奖励信号,可使顶尖视频生成模型的物理合理性提升约6%。
原文摘要 · Abstract (English)
This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical understanding, and often produce implausible videos. The Physics IQ benchmark has shown that visual realism does not imply physics understanding. Yet, intuitive physics understanding has shown to emerge from SSL pretraining on natural videos. In this report, we investigate whether we can leverage SSL-based video world models to improve the physics plausibility of video generative models. In particular, we build ontop of the state-of-the-art video generative model MAGI-1 and couple it with the recently introduced Video Joint Embedding Predictive Architecture 2 (VJEPA-2) to guide the generation process. We show that by leveraging VJEPA-2 as reward signal, we can improve the physics plausibility of state-of-the-art video generative models by ~6%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。