arXiv:2606.27741cs.CV2026-06被引 1

让视频生成模型自己想象运动,更符合物理规律。

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

论文配图:SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
图 1 · 摘自论文原文
  • 用自生成视频替代真实视频训练,打破运动耦合陷阱
  • 提升动作真实性与分离度,尤其在罕见或复杂场景中
  • 适合需要精准控制动作的视频生成应用

视频扩散模型虽显著提升了视觉质量,但生成动作常违背物理规律。我们发现普遍存在“运动纠缠”问题,即相机运动与物体运动被意外耦合。根源在于数据偏差及基于重建的训练设计:模型学习复制原始视频中的粗略运动线索,缺乏建模物理合理运动的动力。为此,提出自想象微调(SIFT)方法,使模型从自身生成的视频中学习,而非直接重构真实视频,从而打破重建捷径。进一步引入运动感知判别监督与渐进式难点重播策略,稳定并加速学习。借助自由生成的文本提示,方法可密集覆盖广泛运动空间,包括难以采集的稀有或精细解耦场景。大量实验表明,该方法显著提升生成视频的物理合理性、运动解耦性与可控性。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure, "motion entanglement", the unintended coupling of independent motion sources, such as camera movement and object motion. We identify that this issue stems from data bias and the reconstruction-based training design of diffusion models. Training on noisy videos that still retain coarse motion cues inadvertently encourages the model to replicate existing motion without an incentive to learn how to model kinematically-grounded motions. To address this, we propose a Self-Imagination Fine-Tuning (SIFT) paradigm, which enables the model to learn from its own generated videos rather than directly reconstructing real ones, breaking the reconstruction shortcut. We further employ motion-aware discriminative supervision and a progressive hard-case replay strategy to stabilize and accelerate learning. By leveraging freely-generated text prompts, our method can densely cover a broad motion space, including rare or finely-disentangled scenarios that would be costly to collect as video data. Extensive experiments demonstrate that our approach substantially improves the physical realism, motion disentanglement, and controllability of generated videos.

视频生成扩散模型运动解耦物理合理性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。