arXiv:2503.00948cs.CV2025-03CVPR被引 15

让视频生成模型更听话、动作更大胆,只需三步解耦控制

Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

  • 用轻量适配器注入文本条件,提升运动可控性
  • 训练免费外推放大动作幅度,显著增强动态范围
  • 分阶段解耦参数,在去噪过程中动态调整运动表现

图像到视频(I2V)生成旨在根据给定图像和条件(如文本)合成视频片段。当前的I2V扩散模型常因运动自由度有限或与文本条件冲突而产生不可控动作。为此,本文提出首个将模型融合技术引入I2V领域的外推与解耦框架。该框架包含三个阶段:(1) 在基础I2V-DM上,通过轻量可学习适配器显式注入文本条件,并微调以提升运动可控性;(2) 提出无需训练的外推策略,逆转微调过程,有效扩大运动动态范围;(3) 将两类运动能力的参数分别解耦并注入基础模型,根据去噪时间步动态调整运动感知参数。大量定性和定量实验表明,该框架在运动可控性与动态幅度上均优于现有方法。

原文摘要 · Abstract (English)

Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance of the images. However, current I2V diffusion models (I2V-DMs) often produce videos with limited motion degrees or exhibit uncontrollable motion that conflicts with the textual condition. To address these limitations, we propose a novel Extrapolating and Decoupling framework, which introduces model merging techniques to the I2V domain for the first time. Specifically, our framework consists of three separate stages: (1) Starting with a base I2V-DM, we explicitly inject the textual condition into the temporal module using a lightweight, learnable adapter and fine-tune the integrated model to improve motion controllability. (2) We introduce a training-free extrapolation strategy to amplify the dynamic range of the motion, effectively reversing the fine-tuning process to enhance the motion degree significantly. (3) With the above two-stage models excelling in motion controllability and degree, we decouple the relevant parameters associated with each type of motion ability and inject them into the base I2V-DM. Since the I2V-DM handles different levels of motion controllability and dynamics at various denoising time steps, we adjust the motion-aware parameters accordingly over time. Extensive qualitative and quantitative experiments have been conducted to demonstrate the superiority of our framework over existing methods.

视频生成扩散模型运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。