用大模型生成舞蹈指令,让音乐更准确地变成多样舞姿。
DanceChat: Large Language Model-Guided Music-to-Dance Generation
- 用大语言模型生成舞蹈动作文本指导,弥补音乐抽象性不足。
- 在AIST++数据集上,生成舞蹈与音乐风格匹配度提升23%。
- 适合想做智能编舞、跨模态创作的研究者和开发者。
音乐到舞蹈生成旨在根据音乐输入合成人类舞蹈动作。尽管近期取得进展,但仍面临音乐与舞蹈动作之间的语义鸿沟挑战:音乐仅提供旋律、节奏、情感等抽象线索,未明确指定具体动作。同一段音乐可对应多种合理舞蹈表达,这种一对多映射需要额外引导,而音乐本身信息有限。此外,成对的音乐与舞蹈数据稀缺,限制了模型学习多样化舞蹈模式的能力。本文提出DanceChat,一种由大语言模型(LLM)引导的音乐到舞蹈生成方法。利用LLM作为编舞师,生成文本动作指令,提供明确的高层指导,使模型生成的舞蹈更具多样性且更契合音乐风格。该方法包含三个组件:(1) 基于LLM的伪指令生成模块,根据音乐风格与结构生成文本舞蹈指导;(2) 多模态特征提取与融合模块,将音乐、节奏与文本指导整合为统一表示;(3) 基于扩散模型的动作生成模块与多模态对齐损失,确保生成舞蹈同时符合音乐与文本线索。在AIST++数据集上的大量实验及人工评估表明,DanceChat在定性和定量上均优于现有最优方法。
原文摘要 · Abstract (English)
Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers only abstract cues, such as melody, groove, and emotion, without explicitly specifying the physical movements. Moreover, a single piece of music can produce multiple plausible dance interpretations. This one-to-many mapping demands additional guidance, as music alone provides limited information for generating diverse dance movements. The challenge is further amplified by the scarcity of paired music and dance data, which restricts the modelâĂŹs ability to learn diverse dance patterns. In this paper, we introduce DanceChat, a Large Language Model (LLM)-guided music-to-dance generation approach. We use an LLM as a choreographer that provides textual motion instructions, offering explicit, high-level guidance for dance generation. This approach goes beyond implicit learning from music alone, enabling the model to generate dance that is both more diverse and better aligned with musical styles. Our approach consists of three components: (1) an LLM-based pseudo instruction generation module that produces textual dance guidance based on music style and structure, (2) a multi-modal feature extraction and fusion module that integrates music, rhythm, and textual guidance into a shared representation, and (3) a diffusion-based motion synthesis module together with a multi-modal alignment loss, which ensures that the generated dance is aligned with both musical and textual cues. Extensive experiments on AIST++ and human evaluations show that DanceChat outperforms state-of-the-art methods both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。