用参考视频提升图像生成视频的运动真实感
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- 从参考视频中检索运动特征,通过上下文感知适配增强生成运动
- 在多个模型和场景下显著提升运动真实性,推理开销几乎为零
- 模块化设计支持零样本迁移,只需更新检索库即可适应新领域
图像到视频生成虽因扩散模型进步而取得显著进展,但生成具有真实运动的视频仍极具挑战性。这源于准确建模运动的复杂性,包括物理约束、物体交互及特定领域的动态行为,难以跨场景泛化。为此,我们提出MotionRAG,一种基于检索增强的框架,通过上下文感知运动适配(CAMA)从相关参考视频中适配运动先验,以提升运动真实感。关键技术包括:(i) 基于检索的流水线,利用视频编码器与专用重采样器提取高层运动特征,提炼语义运动表征;(ii) 采用因果变换器架构实现上下文学习式运动适配;(iii) 基于注意力的运动注入适配器,将转移的运动特征无缝融入预训练视频扩散模型。大量实验表明,该方法在多个领域及多种基础模型上均取得显著提升,且推理阶段计算开销可忽略不计。此外,其模块化设计支持仅通过更新检索数据库即可实现零样本泛化至新领域。本研究通过有效检索与迁移运动先验,提升了视频生成系统的核心能力,促进真实运动动态的合成。
原文摘要 · Abstract (English)
Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling motion, which involves capturing physical constraints, object interactions, and domain-specific dynamics that are not easily generalized across diverse scenarios. To address this, we propose MotionRAG, a retrieval-augmented framework that enhances motion realism by adapting motion priors from relevant reference videos through Context-Aware Motion Adaptation (CAMA). The key technical innovations include: (i) a retrieval-based pipeline extracting high-level motion features using video encoder and specialized resamplers to distill semantic motion representations; (ii) an in-context learning approach for motion adaptation implemented through a causal transformer architecture; (iii) an attention-based motion injection adapter that seamlessly integrates transferred motion features into pretrained video diffusion models. Extensive experiments demonstrate that our method achieves significant improvements across multiple domains and various base models, all with negligible computational overhead during inference. Furthermore, our modular design enables zero-shot generalization to new domains by simply updating the retrieval database without retraining any components. This research enhances the core capability of video generation systems by enabling the effective retrieval and transfer of motion priors, facilitating the synthesis of realistic motion dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。