arXiv:2512.18804cs.CVcs.MM2025-12被引 2

用节拍稳定引导舞蹈生成,避免依赖不靠谱的音乐标签

Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation

  • 构建分层节拍专家混合模型,按60-200BPM分组调度不同节奏专家
  • 在多个数据集上实现最佳舞蹈质量和节奏对齐效果
  • 无需音乐类型标签,适合跨风格真实音乐生成场景

音乐到3D舞蹈生成旨在从音乐中合成逼真且节奏同步的人体舞蹈。现有方法常依赖额外的音乐类型标签以提升生成质量,但这些标签通常噪声大、粒度粗、不可得或不足以捕捉真实音乐的多样性,易导致节奏错位或风格漂移。相比之下,我们观察到节拍(tempo)这一反映音乐节奏与速度的核心属性,在不同数据集和流派中相对稳定,通常介于60至200 BPM之间。基于此,我们提出TempoMoE——一种层次化节拍感知的专家混合模块,增强扩散模型的节奏感知能力。TempoMoE将运动专家按节拍范围分组,结合多尺度节拍专家捕捉短时与长时节奏动态。层级化节奏自适应路由机制根据音乐特征动态选择并融合专家,实现无需人工标注音乐类型下的灵活、精准节奏对齐生成。大量实验表明,TempoMoE在舞蹈质量与节奏对齐方面均达到当前最优表现。

原文摘要 · Abstract (English)

Music to 3D dance generation aims to synthesize realistic and rhythmically synchronized human dance from music. While existing methods often rely on additional genre labels to further improve dance generation, such labels are typically noisy, coarse, unavailable, or insufficient to capture the diversity of real-world music, which can result in rhythm misalignment or stylistic drift. In contrast, we observe that tempo, a core property reflecting musical rhythm and pace, remains relatively consistent across datasets and genres, typically ranging from 60 to 200 BPM. Based on this finding, we propose TempoMoE, a hierarchical tempo-aware Mixture-of-Experts module that enhances the diffusion model and its rhythm perception. TempoMoE organizes motion experts into tempo-structured groups for different tempo ranges, with multi-scale beat experts capturing fine- and long-range rhythmic dynamics. A Hierarchical Rhythm-Adaptive Routing dynamically selects and fuses experts from music features, enabling flexible, rhythm-aligned generation without manual genre labels. Extensive experiments demonstrate that TempoMoE achieves state-of-the-art results in dance quality and rhythm alignment.

舞蹈生成节拍感知专家混合扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。