arXiv:2508.16911cs.GRcs.CV2025-08ICCV被引 8

首个融合文本、音乐与双人舞动的基准数据集,支持智能编舞生成。

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

  • 构建620分钟专业双人舞动捕获数据,同步音乐与万级细粒度描述
  • 提出双任务:文本+音乐生成双人舞,或补全伴舞动作
  • 首次实现文本-音乐-动作三者协同的双人舞蹈生成

我们提出了多模态双人舞(MDD)数据集,一个用于文本控制和音乐条件化3D双人舞动生成的多样化多模态基准。数据集包含由专业舞者表演的620分钟高质量动作捕捉数据,与音乐同步,并配有超过10,000条细粒度自然语言描述。标注内容涵盖丰富的动作词汇,包括空间关系、身体运动与节奏特征,使MDD成为首个无缝整合人类动作、音乐与文本的双人舞生成数据集。我们引入两个新任务:(1) 文本到双人舞——给定音乐与文本提示,生成领舞与伴舞的动作;(2) 文本到舞蹈伴奏——给定音乐、文本提示及领舞动作,生成与文本一致且连贯的伴舞动作。我们在两个任务上提供了基线评估,以支持未来研究。

原文摘要 · Abstract (English)

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and detailed with over 10K fine-grained natural language descriptions. The annotations capture a rich movement vocabulary, detailing spatial relationships, body movements, and rhythm, making MDD the first dataset to seamlessly integrate human motions, music, and text for duet dance generation. We introduce two novel tasks supported by our dataset: (1) Text-to-Duet, where given music and a textual prompt, both the leader and follower dance motion are generated (2) Text-to-Dance Accompaniment, where given music, textual prompt, and the leader's motion, the follower's motion is generated in a cohesive, text-aligned manner. We include baseline evaluations on both tasks to support future research.

双人舞生成多模态数据集文本控制动作合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。