首个实时可控运动的自回归视频生成模型,支持多样运动控制且延迟极低。
Real-Time Motion-Controllable Autoregressive Video Diffusion
- 基于强化学习优化自回归视频扩散,实现运动可控生成。
- 仅用1.3B参数,少步数生成下保持高画质与精准运动对齐。
- 适合需要低延迟、多样化运动控制的实时视频生成场景。
实时可控运动的视频生成仍面临双向扩散模型固有的延迟及有效自回归(AR)方法缺失的挑战。现有AR视频扩散模型多限于简单控制信号或文本到视频生成,且在少步数生成时易出现质量下降和运动伪影。为此,我们提出AR-Drag,首个基于强化学习的少步数自回归视频扩散模型,支持实时图像到视频生成并实现多样运动控制。首先微调基础图像到视频模型以支持基础运动控制,再通过轨迹奖励模型的强化学习进一步优化。设计采用自滚动机制保持马尔可夫性,并通过选择性引入去噪步骤的随机性加速训练。大量实验表明,AR-Drag在视觉保真度和运动对齐上表现优异,相比最先进可控视频扩散模型显著降低延迟,仅使用1.3B参数。更多可视化见项目页:https://kesenzhao.github.io/AR-Drag.github.io/。
原文摘要 · Abstract (English)
Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the Markov property through a Self-Rollout mechanism and accelerates training by selectively introducing stochasticity in denoising steps. Extensive experiments demonstrate that AR-Drag achieves high visual fidelity and precise motion alignment, significantly reducing latency compared with state-of-the-art motion-controllable VDMs, while using only 1.3B parameters. Additional visualizations can be found on our project page: https://kesenzhao.github.io/AR-Drag.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。