让舞蹈生成精准匹配音乐风格,支持文本控制风格与节奏同步。
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
- 用文本映射生成风格化控制信号,实现风格可控的舞蹈生成。
- 在FineDance和AIST++数据集上优于现有最先进方法。
- 适合需要精确风格控制的舞蹈动画、虚拟演出场景。
音乐驱动的舞蹈生成极具挑战性,需严格遵循特定音乐风格的编舞,同时保证动作物理真实且与音乐节拍和节奏精确同步。尽管已有进展,多数方法仍难以准确传达生成舞蹈的特定风格特征。为此,我们提出一种基于扩散模型的框架,以音乐和描述性文本为条件,实现风格特定的3D全身舞蹈生成。为有效融入风格信息,我们设计了一种基于文本的控制机制,将输入提示(显式风格标签或自由描述文本)映射为风格特定的控制信号,实现精确可控的风格一致舞蹈生成。此外,为增强音乐与文本条件间的对齐,我们利用音乐基础模型的特征,促进连贯且语义一致的舞蹈合成。最后,为平衡提取文本-风格信息与保持高质量生成之间的目标,我们提出一种新型多任务优化策略,有效协调物理真实性、空间精度与文本分类等相互竞争因素,显著提升生成序列的整体质量。在FineDance和AIST++数据集上的大量实验结果表明,GCDance优于现有最先进方法。
原文摘要 · Abstract (English)
Music-driven dance generation is a challenging task as it requires strict adherence to genre-specific choreography while ensuring physically realistic and precisely synchronized dance sequences with the music's beats and rhythm. Although significant progress has been made in music-conditioned dance generation, most existing methods struggle to convey specific stylistic attributes in generated dance. To bridge this gap, we propose a diffusion-based framework for genre-specific 3D full-body dance generation, conditioned on both music and descriptive text. To effectively incorporate genre information, we develop a text-based control mechanism that maps input prompts, either explicit genre labels or free-form descriptive text, into genre-specific control signals, enabling precise and controllable text-guided generation of genre-consistent dance motions. Furthermore, to enhance the alignment between music and textual conditions, we leverage the features of a music foundation model, facilitating coherent and semantically aligned dance synthesis. Last, to balance the objectives of extracting text-genre information and maintaining high-quality generation results, we propose a novel multi-task optimization strategy. This effectively balances competing factors such as physical realism, spatial accuracy, and text classification, significantly improving the overall quality of the generated sequences. Extensive experimental results obtained on the FineDance and AIST++ datasets demonstrate the superiority of GCDance over the existing state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。