arXiv:2604.17005cs.CVcs.SD2026-04

让舞蹈生成听懂自然语言指令,实现精准语义控制。

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

论文配图:TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
图 1 · 摘自论文原文
  • 以动作作为桥梁,统一音乐-动作与文本-动作数据的语义空间。
  • 在不依赖人工标注三元组的情况下,实现高质量舞蹈生成与文本控制。
  • 适合需要自然语言引导舞蹈创作的研究者与开发者。

现有音乐驱动舞蹈生成方法虽具备高真实感和良好的音频-动作对齐效果,但普遍缺乏语义可控性,难以通过自然语言描述引导特定动作。其核心瓶颈在于缺少大规模联合对齐音乐、文本与动作的标注数据集。为此,我们提出TeMuDance框架,无需任何手动标注的音乐-文本-动作三元组数据,即可实现文本条件下的舞蹈生成控制。TeMuDance采用以动作为中心的桥梁范式,利用动作作为共享语义锚点,将分离的音乐-动作与文本-动作数据集映射到统一嵌入空间,支持跨模态检索缺失模态以进行端到端训练。在此基础上,构建轻量级文本控制分支,在冻结的音乐-动作扩散主干上训练,保持节奏准确性的同时实现细粒度语义引导。为抑制检索监督中的噪声,设计基于置信度筛选的双流微调策略。此外,提出一种任务对齐的评估指标,量化文本提示是否在音乐条件下引发预期的运动特征。大量实验表明,TeMuDance在保持舞蹈质量竞争力的同时,显著优于现有方法的文本条件控制能力。

原文摘要 · Abstract (English)

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements through natural language descriptions. This limitation primarily stems from the absence of large-scale datasets that jointly align music, text, and motion for supervised learning of text-conditioned control. To address this challenge, we propose TeMuDance, a framework that enables text-based control for music-conditioned dance generation without requiring any manually annotated music-text-motion triplet dataset. TeMuDance introduces a motion-centred bridging paradigm that leverages motion as a shared semantic anchor to align disjoint music-dance and text-motion datasets within a unified embedding space, enabling cross-modal retrieval of missing modalities for end-to-end training. A lightweight text control branch is then trained on top of a frozen music-to-dance diffusion backbone, preserving rhythmic fidelity while enabling fine-grained semantic guidance. To further suppress noise inherent in the retrieved supervision, we design a dual-stream fine-tuning strategy with confidence-based filtering. We also propose a novel task-aligned metric that quantifies whether textual prompts induce the intended kinematic attributes under music conditioning. Extensive experiments demonstrate that TeMuDance achieves competitive dance quality while substantially improving text-conditioned control over existing methods.

舞蹈生成文本控制跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。