提升视频生成模型的运动一致性,解决微调后动作失真问题。
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
- 用像素运动动态设计奖励机制,捕捉瞬时与长期运动一致性。
- 提出可扩展的对抗性混合微调框架,显著加快收敛并防止奖励滥用。
- 适合关注视频生成质量、尤其是运动连贯性的研究者和开发者。
图像条件视频扩散模型虽具备出色的视觉真实感,但在微调后常出现运动保真度下降的问题,如运动动态减弱或长期时间连贯性退化。本文研究了训练后视频扩散模型中的运动对齐问题。为此,我们基于像素通量动态设计了像素运动奖励,同时捕捉瞬时与长期的运动一致性。进一步提出了平滑混合微调(SHIFT)框架,该框架统一了监督微调与优势加权微调,通过新型对抗性优势机制,提升了收敛速度并缓解了奖励劫持问题。实验表明,该方法能有效解决现代视频扩散模型在监督微调中出现的动态程度坍塌问题。项目主页:https://xiye20.github.io/projects/SHIFT/
原文摘要 · Abstract (English)
Image-conditioned video diffusion models achieve impressive visual realism but often suffer from weakened motion fidelity, e.g., reduced motion dynamics or degraded long-term temporal coherence, especially after fine-tuning. We study motion alignment in video diffusion models post-training. To address this, we introduce pixel-motion rewards based on pixel flux dynamics, capturing both instantaneous and long-term motion consistency. We further propose \underline{S}mooth \underline{H}ybr\underline{i}d \underline{F}ine-\underline{t}uning (SHIFT), a scalable reward-driven framework that unifies supervised fine-tuning and advantage-weighted fine-tuning. Benefiting from novel adversarial advantages, SHIFT improves convergence speed and mitigates reward hacking. Experiments show that our approach efficiently resolves dynamic-degree collapse in modern video diffusion models supervised fine-tuning. Project page: https://xiye20.github.io/projects/SHIFT/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。