arXiv:2512.18181cs.CV2025-12International Conf…被引 14

MACE-Dance 用分层专家模型,让音乐自动生成动作与视觉一致的舞蹈视频。

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

  • 分两阶段:先生成符合音乐的动作,再根据动作和参考图合成视频。
  • 在公开数据集上实现当前最优的舞蹈动作真实感和画面一致性。
  • 适合想做音乐驱动视频生成的研究者和开发者。

随着在线舞蹈平台兴起和AIGC技术快速发展,音乐驱动舞蹈生成成为重要研究方向。尽管在3D舞蹈生成、姿态驱动图像动画和音频驱动人脸合成等领域已有进展,现有方法难以直接适用于该任务,且多数仍无法同时保证高质量视觉外观与真实人体运动。为此,我们提出MACE-Dance,一种基于级联混合专家(MoE)的音乐驱动舞蹈视频生成框架。运动专家通过扩散模型结合BiMamba-Transformer架构与无引导训练(GFT)策略,生成具有运动合理性与艺术表现力的3D动作;外观专家则在动作与参考图像条件下进行视频合成,保持身份一致性与时空连贯性。此外,我们构建了一个大规模多样化的数据集,并设计了运动-外观联合评估协议。基于该协议,MACE-Dance在多项指标上达到当前最优性能。代码已开源。

原文摘要 · Abstract (English)

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domains such as music-driven 3D dance generation, pose-driven image animation, and audio-driven talking-head synthesis, existing methods cannot be directly adapted to this task. Moreover, the limited studies in this area still struggle to jointly achieve high-quality visual appearance and realistic human motion. Accordingly, we present MACE-Dance, a music-driven dance video generation framework with cascaded Mixture-of-Experts (MoE). The Motion Expert performs music-to-3D motion generation while enforcing kinematic plausibility and artistic expressiveness, whereas the Appearance Expert carries out motion- and reference-conditioned video synthesis, preserving visual identity with spatiotemporal coherence. Specifically, the Motion Expert adopts a diffusion model with a BiMamba-Transformer hybrid architecture and a Guidance-Free Training (GFT) strategy, achieving state-of-the-art (SOTA) performance in 3D dance generation. The Appearance Expert employs a decoupled kinematic-aesthetic fine-tuning strategy, achieving state-of-the-art (SOTA) performance in pose-driven image animation. To better benchmark this task, we curate a large-scale and diverse dataset and design a motion-appearance evaluation protocol. Based on this protocol, MACE-Dance also achieves state-of-the-art performance. Code is available at https://github.com/AMAP-ML/MACE-Dance.

舞蹈生成扩散模型多模态生成专家系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。