arXiv:2411.01805cs.SDcs.MM2024-11NeurIPS被引 7

提出统一生成舞步与音乐的扩散模型,实现长时序同步

MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence

  • 用双向节奏变分自编码器对齐动作与音乐潜在表示
  • 在长达128秒的序列上保持节奏同步,生成质量优于现有方法
  • 适合舞蹈生成、跨模态创作及长序列内容设计研究者

动作到音乐与音乐到动作的研究长期分开进行,二者交互体现人类高级智能,建立统一关联至关重要。然而至今尚无工作联合探索两者间的模态对齐。为此,我们提出新框架MoMu-Diffusion,用于长时序同步的动作-音乐生成。首先,为降低长序列带来的巨大计算开销,提出双向对比节奏变分自编码器(BiCoR-VAE),提取动作与音乐输入的模态对齐潜在表征。随后,基于对齐的潜在空间,引入多模态Transformer扩散模型与交叉引导采样策略,支持跨模态、多模态及可变长度生成任务。大量实验表明,MoMu-Diffusion在定性与定量上均超越近期最先进方法,能合成真实、多样、长时序且节拍匹配的音乐或动作序列。生成样本与代码已公开于https://momu-diffusion.github.io/

原文摘要 · Abstract (English)

Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motion and music is a reflection of advanced human intelligence, and establishing a unified relationship between them is particularly important. However, to date, there has been no work that considers them jointly to explore the modality alignment within. To bridge this gap, we propose a novel framework, termed MoMu-Diffusion, for long-term and synchronous motion-music generation. Firstly, to mitigate the huge computational costs raised by long sequences, we propose a novel Bidirectional Contrastive Rhythmic Variational Auto-Encoder (BiCoR-VAE) that extracts the modality-aligned latent representations for both motion and music inputs. Subsequently, leveraging the aligned latent spaces, we introduce a multi-modal Transformer-based diffusion model and a cross-guidance sampling strategy to enable various generation tasks, including cross-modal, multi-modal, and variable-length generation. Extensive experiments demonstrate that MoMu-Diffusion surpasses recent state-of-the-art methods both qualitatively and quantitatively, and can synthesize realistic, diverse, long-term, and beat-matched music or motion sequences. The generated samples and codes are available at https://momu-diffusion.github.io/

动作生成音乐生成扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。