arXiv:2603.08023cs.CVcs.AI2026-03中稿 · WACV 2026被引 1

用Mamba模型生成更符合节奏的舞蹈动作,解决传统方法对音乐律动捕捉不足的问题。

Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model

  • 采用Mamba替代Transformer,更好处理长序列舞蹈的时序依赖
  • 引入高斯化节拍表示,显式引导舞蹈动作与音乐节拍对齐
  • 在AIST++和FineDance数据集上实现长短舞均稳定生成,优于现有方法

舞蹈是体现情感与交流的人体运动形式,在音乐、虚拟现实和内容创作中具有重要应用。现有舞蹈生成方法难以充分捕捉舞蹈固有的顺序性、节奏性和音乐同步特性。本文提出基于Mamba的扩散模型MambaDance,将擅长处理长序列自回归结构的Mamba引入双阶段扩散架构,替代传统的Transformer。同时,为强化音乐节拍在编舞中的作用,设计了一种高斯基节拍表示,显式指导舞蹈序列解码。在AIST++和FineDance数据集上,针对不同长度序列的实验表明,该方法能持续生成合理且具备关键特征的舞蹈动作,从短到长舞蹈均表现优异,显著优于先前方法。更多定性结果与演示视频见https://vision3d-lab.github.io/mambadance。

原文摘要 · Abstract (English)

Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music-synchronized characteristics of dance. In this paper, we propose \emph{MambaDance}, a new dance generation approach that leverages a Mamba-based diffusion model. Mamba, well-suited to handling long and autoregressive sequences, is integrated into our two-stage diffusion architecture, substituting off-the-shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian-based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at \small{https://vision3d-lab.github.io/mambadance}.

舞蹈生成Mamba模型扩散模型节奏对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。