用频域物理先验提升视频生成运动合理性,不改模型架构
Physics-Guided Motion Loss for Video Generation Model
- 通过频域分解刚性运动,设计轻量级谱损失
- 在OpenVID-1M上动作识别提升11%,视觉质量不变
- 适合关注视频运动真实性的生成模型开发者
当前视频扩散模型虽视觉逼真,但常违背基本物理规律,产生如橡胶变形、物体运动不一致等细微瑕疵。本文提出一种频域物理先验,无需修改模型结构即可提升运动合理性。方法将常见刚性运动(平移、旋转、缩放)分解为轻量级谱损失,仅需2.7%的频率系数即可保留97%以上的谱能量。应用于Open-Sora、MVDIT和Hunyuan模型,在OpenVID-1M数据集上平均提升动作识别准确率约11%(相对),同时保持视觉质量。用户实验显示,74%-83%偏好本方法生成的视频。该方法还可降低22%-37%的形变误差,并提升时序一致性得分。结果表明,简单全局谱特征可作为视频扩散模型中物理合理运动的有效即插即用正则化器。
原文摘要 · Abstract (English)
Current video diffusion models generate visually compelling content but often violate basic laws of physics, producing subtle artifacts like rubber-sheet deformations and inconsistent object motion. We introduce a frequency-domain physics prior that improves motion plausibility without modifying model architectures. Our method decomposes common rigid motions (translation, rotation, scaling) into lightweight spectral losses, requiring only 2.7% of frequency coefficients while preserving 97%+ of spectral energy. Applied to Open-Sora, MVDIT, and Hunyuan, our approach improves both motion accuracy and action recognition by ~11% on average on OpenVID-1M (relative), while maintaining visual quality. User studies show 74--83% preference for our physics-enhanced videos. It also reduces warping error by 22--37% (depending on the backbone) and improves temporal consistency scores. These results indicate that simple, global spectral cues are an effective drop-in regularizer for physically plausible motion in video diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。