用统一编码提升舞蹈连贯性,实现高质量音乐生成3D舞蹈
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
- 通过运动学动力学约束的量化编码构建潜在表示
- 在FineDance数据集上达到当前最优性能
- 适合虚拟现实与创意内容生成领域研究者
音乐到舞蹈生成是编舞、虚拟现实和创意内容生成交叉领域的关键挑战任务。现有方法在保持舞蹈连贯性方面存在明显局限。为此,我们提出MatchDance框架,通过构建潜在表示来增强舞蹈连贯性。该框架采用两阶段设计:(1) 基于运动学-动力学约束的量化阶段(KDQS),利用有限标量量化(FSQ)将舞蹈动作编码为潜在表示,并高保真重建;(2) 混合音乐到舞蹈生成阶段(HMDGS),采用Mamba-Transformer混合架构将音乐映射至潜在空间,再通过KDQS解码器生成3D舞蹈动作。此外,引入音乐-舞蹈检索框架和综合评估指标。在FineDance数据集上的大量实验表明,该方法达到当前最佳性能。
原文摘要 · Abstract (English)
Music-to-dance generation represents a challenging yet pivotal task at the intersection of choreography, virtual reality, and creative content generation. Despite its significance, existing methods face substantial limitation in achieving choreographic consistency. To address the challenge, we propose MatchDance, a novel framework for music-to-dance generation that constructs a latent representation to enhance choreographic consistency. MatchDance employs a two-stage design: (1) a Kinematic-Dynamic-based Quantization Stage (KDQS), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) with kinematic-dynamic constraints and reconstructs them with high fidelity, and (2) a Hybrid Music-to-Dance Generation Stage(HMDGS), which uses a Mamba-Transformer hybrid architecture to map music into the latent representation, followed by the KDQS decoder to generate 3D dance motions. Additionally, a music-dance retrieval framework and comprehensive metrics are introduced for evaluation. Extensive experiments on the FineDance dataset demonstrate state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。