用海量数据和多模态控制生成逼真3D舞蹈,支持音乐+文本/动作轨迹自由组合
OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
- 构建100小时跨14种风格的多模态舞蹈数据集,含3D动作、音乐、关键点等标注
- 提出统一建模框架,可基于音乐+任意组合的文本、关键点或轨迹生成舞蹈
- 生成动作自然流畅,物理接触真实,适合动画、游戏、虚拟演出等场景
音乐驱动的3D舞蹈生成具有巨大创意潜力,但实际应用需具备多样化和多模态控制能力。由于舞蹈动作高度动态且复杂,涵盖多种风格与流派,生成过程需满足音乐之外的多种条件(如空间轨迹、关键帧手势或风格描述)。然而,缺乏大规模、丰富标注的数据集严重制约了该领域发展。本文构建了OpenDanceSet,一个包含超过100小时、覆盖14种风格和147名舞者的大型人体舞蹈数据集。每个样本均配有丰富的标注信息,包括3D动作、配对音乐、2D关键点、运动轨迹及专家标注的文本描述,支持鲁棒的跨模态学习。同时,我们提出OpenDanceNet,一种统一的掩码建模框架,包含解耦自编码器与多模态联合预测Transformer,可实现基于音乐以及文本、关键点或轨迹任意组合的可控生成。大量实验表明,本方法在保持高保真度的同时,展现出强多样性与真实的物理接触行为,并能灵活控制空间与风格条件。
原文摘要 · Abstract (English)
Music-driven 3D dance generation offers significant creative potential, yet practical applications demand versatile and multimodal control. As the highly dynamic and complex human motion covering various styles and genres, dance generation requires satisfying diverse conditions beyond just music (e.g., spatial trajectories, keyframe gestures, or style descriptions). However, the absence of a large-scale and richly annotated dataset severely hinders progress. In this paper, we build OpenDanceSet, an extensive human dance dataset comprising over 100 hours across 14 genres and 147 subjects. Each sample has rich annotations to facilitate robust cross-modal learning: 3D motion, paired music, 2D keypoints, trajectories, and expert-annotated text descriptions. Furthermore, we propose OpenDanceNet, a unified masked modeling framework for controllable dance generation, including a disentangled auto-encoder and a multimodal joint-prediction Transformer. OpenDanceNet supports generation conditioned on music and arbitrary combinations of text, keypoints, or trajectories. Comprehensive experiments demonstrate that our work achieves high-fidelity synthesis with strong diversity and realistic physical contacts, while also offering flexible control over spatial and stylistic conditions. Project Page: https://open-dance.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。