arXiv:2505.23656cs.CV2025-05NeurIPS被引 87

通过关系对齐,让视频生成模型学会更符合物理规律的运动逻辑。

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

  • 用令牌级关系对齐,将基础模型的物理理解能力迁移到生成模型中。
  • 在基准测试上显著提升基线模型的物理常识表现,生成更合理的动态画面。
  • 适合需要高物理真实性视频生成的研究者与开发者使用。

近年来,文本到视频(T2V)扩散模型已实现高质量、逼真的视频合成。然而,现有T2V模型常因缺乏准确理解物理规律的能力,难以生成符合物理常识的内容。我们发现,尽管T2V模型内部表征具备一定物理理解潜力,但远落后于最新的视频自监督学习方法。为此,我们提出VideoREPA框架,通过对齐令牌级关系,将视频理解基础模型的物理理解能力蒸馏至T2V模型中,弥合物理理解差距,实现更符合物理规律的生成。具体而言,引入了令牌关系蒸馏(TRD)损失,利用时空对齐提供适用于微调强大预训练T2V模型的软性指导,这是对先前表示对齐(REPA)方法的重要突破。据我们所知,VideoREPA是首个专为微调T2V模型设计、并用于注入物理知识的REPA方法。实证评估表明,VideoREPA显著提升了基线模型CogVideoX的物理常识能力,在相关基准上取得显著改进,并展现出强一致性物理生成能力。更多视频结果见 https://videorepa.github.io/。

原文摘要 · Abstract (English)

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the representations within T2V models possess some capacity for physics understanding, they lag significantly behind those from recent video self-supervised learning methods. To this end, we propose a novel framework called VideoREPA, which distills physics understanding capability from video understanding foundation models into T2V models by aligning token-level relations. This closes the physics understanding gap and enable more physics-plausible generation. Specifically, we introduce the Token Relation Distillation (TRD) loss, leveraging spatio-temporal alignment to provide soft guidance suitable for finetuning powerful pre-trained T2V models, a critical departure from prior representation alignment (REPA) methods. To our knowledge, VideoREPA is the first REPA method designed for finetuning T2V models and specifically for injecting physical knowledge. Empirical evaluations show that VideoREPA substantially enhances the physics commonsense of baseline method, CogVideoX, achieving significant improvement on relevant benchmarks and demonstrating a strong capacity for generating videos consistent with intuitive physics. More video results are available at https://videorepa.github.io/.

视频生成物理理解蒸馏扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。