arXiv:2502.20176cs.SDcs.GR2025-02中稿 · the Audio Imaginat…被引 7

用音乐大模型生成更真实、贴合音乐的全身舞蹈动作

DGFM: Full Body Dance Generation Driven by Music Foundation Models

  • 结合音乐大模型与手工特征,提升音乐理解能力
  • 生成动作与音乐匹配度最高,真实感强于现有方法
  • 适合需要音乐驱动舞蹈生成的研究与创作场景

在音乐驱动的舞蹈动作生成中,现有方法多依赖手工特征,忽视了音乐大模型对跨模态内容生成的深远影响。为此,我们提出一种基于扩散模型的方法,通过融合音乐大模型提取的高层语义特征与手工特征,生成受文本和音乐共同控制的舞蹈动作。该方法有效结合了高层语义信息与低层时序细节,显著提升了模型对音乐特征的理解能力。为验证其优势,我们与四种音乐大模型及两组手工特征进行了对比。结果表明,本方法生成的舞蹈序列最为真实,且与输入音乐的匹配度最优。

原文摘要 · Abstract (English)

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based method that generates dance movements conditioned on text and music. Our approach extracts music features by combining high-level features obtained by music foundation model with hand-crafted features, thereby enhancing the quality of generated dance sequences. This method effectively leverages the advantages of high-level semantic information and low-level temporal details to improve the model's capability in music feature understanding. To show the merits of the proposed method, we compare it with four music foundation models and two sets of hand-crafted music features. The results demonstrate that our method obtains the most realistic dance sequences and achieves the best match with the input music.

舞蹈生成音乐理解扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。