用海量舞蹈数据和新框架,让视频生成更懂音乐节奏
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

- 构建300万片段的舞蹈数据集,融合动作与音乐信息
- 在400小时数据上实现音乐驱动的高质量舞蹈生成
- 适合做多模态视频生成、舞蹈动画研究者使用
音乐驱动的舞蹈视频生成旨在合成与音乐时间对齐且视觉质量高的表达性人体运动。尽管已有进展,现有方法仍面临两大瓶颈:缺乏大规模高质量舞蹈视频数据集,以及缺乏将音乐作为互补条件信号整合进视频生成基础模型的系统性框架。为此,我们提出CIPE-Dance,一个基于互联网的大规模舞蹈视频数据集,通过渐进式专家流程构建,包含编舞导向的文本标注。据我们所知,CIPE-Dance是目前最大的舞蹈视频生成数据集,涵盖300,000个高质量片段,总时长400小时,覆盖多样舞者、场景与舞种。我们进一步提出OmniDance,一种无需牺牲原始可控性或视觉保真度即可将音乐融入TI2V基础模型的框架级方案。受文本(低频语义)与音乐(高频时序动态)互补性的启发,OmniDance设计了深度感知专业化架构、锚定式由易到难课程学习策略及模态专用的时间依赖CFG策略,实现TI2V、MI2V与MTI2V的统一生成。在CIPE-Dance上的大量实验表明,OmniDance在三项任务中均达到当前最优性能,并展现出强健的多模态融合能力。项目地址:https://github.com/AMAP-ML/OmniDance。
原文摘要 · Abstract (English)
Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models. To address these limitations, we introduce CIPE-Dance, a large-scale Internet-sourced dance video dataset with choreography-informed text annotations, constructed via a progressive expert pipeline. To the best of our knowledge, CIPE-Dance is the largest dataset for dance video generation to date, comprising 300k high-quality clips over 400 hours and covering diverse dancers, environments, and dance genres. We further propose OmniDance, a framework-level recipe for integrating music into a TI2V foundation model without sacrificing its original controllability or visual fidelity. Motivated by the complementary roles of text as low-frequency semantics and music as high-frequency temporal dynamics, OmniDance co-designs a depth-aware specialization architecture, an anchored easy-to-hard curriculum learning strategy, and a modality-specialized time-dependent CFG strategy, enabling unified TI2V, MI2V, and MTI2V generation. Extensive experiments on CIPE-Dance demonstrate that OmniDance achieves state-of-the-art performance across all three tasks and exhibits robust multimodal integration capability. Project is available at https://github.com/AMAP-ML/OmniDance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。