arXiv:2601.03323cs.GRcs.CV2026-01

用节奏和文字生成连贯舞蹈,支持长序列自回归输出

Listen to Rhythm, Choose Movements: Autoregressive Multimodal Dance Generation via Diffusion and Mamba with Decoupled Dance Dataset

  • 分离舞蹈数据的音轨、动作与文本描述,构建解耦数据集
  • 结合扩散模型与Mamba模块,实现多模态引导的长序列舞蹈生成
  • 适合对舞蹈生成多样性和时序连贯性有要求的研究者

生成模型与序列学习的进步推动了舞蹈动作生成研究,但现有方法仍存在语义控制粗略、长序列连贯性差的问题。本文提出听节拍选动作(LRCM)框架,支持多模态输入与自回归舞蹈生成。我们对舞蹈数据采用特征解耦范式,推广至Motorica Dance数据集,将动作捕捉数据、音频节拍以及专业标注的全局与局部文本描述分离。扩散架构融合音频潜空间Conformer与文本潜空间跨注意力机制,并引入运动时间Mamba模块(MTMM),实现平滑的长时序自回归合成。实验表明,LRCM在功能表现与定量指标上均表现出色,尤其在多模态输入与长序列生成场景中展现出显著潜力。项目页面见 https://oranduanstudy.github.io/LRCM/。

原文摘要 · Abstract (English)

Advances in generative models and sequence learning have greatly promoted research in dance motion generation, yet current methods still suffer from coarse semantic control and poor coherence in long sequences. In this work, we present Listen to Rhythm, Choose Movements (LRCM), a multimodal-guided diffusion framework supporting both diverse input modalities and autoregressive dance motion generation. We explore a feature decoupling paradigm for dance datasets and generalize it to the Motorica Dance dataset, separating motion capture data, audio rhythm, and professionally annotated global and local text descriptions. Our diffusion architecture integrates an audio-latent Conformer and a text-latent Cross-Conformer, and incorporates a Motion Temporal Mamba Module (MTMM) to enable smooth, long-duration autoregressive synthesis. Experimental results indicate that LRCM delivers strong performance in both functional capability and quantitative metrics, demonstrating notable potential in multimodal input scenarios and extended sequence generation. The project page is available at https://oranduanstudy.github.io/LRCM/.

舞蹈生成多模态扩散模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。