让视频生成模型学会真实运动,提升物理合理性
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
- 从预训练视频编码器中分离出独立的运动特征空间
- 通过光流预测优化运动表征,使生成视频更符合物理规律
- 适合关注视频时序连贯性与真实感的生成研究者
文本到视频扩散模型虽能生成高质量视频,但常缺乏时间连贯性和物理合理性。主要原因在于模型对复杂运动理解不足。现有方法尝试对齐扩散模型特征与预训练视频编码器特征,但这些编码器将外观与动态混合,限制了对齐效果。本文提出一种以运动为中心的对齐框架,从预训练视频编码器中学习解耦的运动子空间,并通过优化其对真实光流的预测能力,确保该空间捕捉真实运动动态。随后将文本到视频扩散模型的潜在特征对齐至该运动子空间,使生成模型内化运动知识,从而生成更具合理性的视频。在VideoPhy、VideoPhy2、VBench和VBench-2.0上的实证评估及用户研究均表明,该方法显著提升视频的物理常识性,同时保持对文本提示的忠实度。
原文摘要 · Abstract (English)
Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by aligning diffusion model features with those from pretrained video encoders. However, these encoders mix video appearance and dynamics into entangled features, limiting the benefit of such alignment. In this paper, we propose a motion-centric alignment framework that learns a disentangled motion subspace from a pretrained video encoder. This subspace is optimized to predict ground-truth optical flow, ensuring it captures true motion dynamics. We then align the latent features of a text-to-video diffusion model to this new subspace, enabling the generative model to internalize motion knowledge and generate more plausible videos. Our method improves the physical commonsense in a state-of-the-art video diffusion model, while preserving adherence to textual prompts, as evidenced by empirical evaluations on VideoPhy, VideoPhy2, VBench, and VBench-2.0, along with a user study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。