用检索+扩散模型生成长时序音乐舞蹈,更连贯更配乐。
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
- 用对比学习对齐音乐与舞蹈表征,无配对数据也能对齐语义。
- 优化动作图实现高效检索拼接,保证长序列动作真实连贯。
- 结合原始音乐与特征的扩散模型,提升动作质量与节奏同步。
生成长期、连贯且真实的音乐驱动舞蹈序列仍是人体运动合成中的难题。现有方法存在明显局限:动作图方法依赖固定模板库,限制创作自由;扩散模型虽能生成新动作,但常缺乏时间连贯性和音乐对齐。为此,我们提出MotionRAG-Diff,一种融合检索增强生成(RAG)与扩散模型精炼的混合框架,可为任意长音乐输入生成高质量、音乐同步的舞蹈序列。本方法提出三项核心创新:(1) 采用跨模态对比学习架构,在共享隐空间中对齐异构的音乐与舞蹈表示,实现无配对数据下的语义对应;(2) 优化动作图系统,实现动作片段的高效检索与无缝拼接,确保长序列的动作真实性和时间连贯性;(3) 设计多条件扩散模型,联合依赖原始音乐信号与对比特征,提升动作质量与全局节奏同步。大量实验表明,MotionRAG-Diff在动作质量、多样性与音乐-动作同步准确率上均达到当前最优水平。该工作通过融合检索模板保真度与扩散模型创造力,建立音乐驱动舞蹈生成的新范式。
原文摘要 · Abstract (English)
Generating long-term, coherent, and realistic music-conditioned dance sequences remains a challenging task in human motion synthesis. Existing approaches exhibit critical limitations: motion graph methods rely on fixed template libraries, restricting creative generation; diffusion models, while capable of producing novel motions, often lack temporal coherence and musical alignment. To address these challenges, we propose $\textbf{MotionRAG-Diff}$, a hybrid framework that integrates Retrieval-Augmented Generation (RAG) with diffusion-based refinement to enable high-quality, musically coherent dance generation for arbitrary long-term music inputs. Our method introduces three core innovations: (1) A cross-modal contrastive learning architecture that aligns heterogeneous music and dance representations in a shared latent space, establishing unsupervised semantic correspondence without paired data; (2) An optimized motion graph system for efficient retrieval and seamless concatenation of motion segments, ensuring realism and temporal coherence across long sequences; (3) A multi-condition diffusion model that jointly conditions on raw music signals and contrastive features to enhance motion quality and global synchronization. Extensive experiments demonstrate that MotionRAG-Diff achieves state-of-the-art performance in motion quality, diversity, and music-motion synchronization accuracy. This work establishes a new paradigm for music-driven dance generation by synergizing retrieval-based template fidelity with diffusion-based creative enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。