用100小时真实舞蹈数据训练大模型,生成更自然的3D跳舞动作。
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
- 自动从单目视频重建高保真3D舞蹈,加入足部约束提升物理合理性。
- 模型支持不同音乐节奏,生成动作流畅且符合音乐节拍。
- 适合做真实场景舞蹈生成,尤其对音乐风格多变的创作有帮助。
现有3D舞蹈生成方法在受控环境下表现良好,但在真实场景中泛化能力差。当面对未见过的音乐时,常生成无结构或物理上不合理的动作,主要因音乐-舞蹈数据有限且模型容量不足。本文通过扩大数据与模型设计来推动可泛化的3D舞蹈生成。首先,在数据方面,开发全自动管道,从单目视频重建高质量3D舞蹈动作。为消除现有方法中的物理伪影,引入基于足部接触与几何约束的足部修复扩散模型(FRDM),在保持运动连贯性与表现力的同时确保物理合理性,构建出总时长达100.69小时的多样化、高质量多模态3D舞蹈数据集。其次,在模型设计上,提出基于LLaMA的可扩展架构Choreographic LLaMA(ChoreoLLaMA)。为增强对陌生音乐的鲁棒性,集成检索增强生成(RAG)模块,注入参考舞蹈作为提示;同时设计慢/快节奏混合专家(MoE)模块,使模型能平滑适应不同音乐速度。大量实验表明,该方法在多种舞蹈风格下均优于现有方法,实现了定性和定量上的提升,标志着向可扩展的真实世界3D舞蹈生成迈进一步。代码、模型与数据将公开。
原文摘要 · Abstract (English)
Although existing 3D dance generation methods perform well in controlled scenarios, they often struggle to generalize in the wild. When conditioned on unseen music, existing methods often produce unstructured or physically implausible dance, largely due to limited music-to-dance data and restricted model capacity. This work aims to push the frontier of generalizable 3D dance generation by scaling up both data and model design. (1) On the data side, we develop a fully automated pipeline that reconstructs high-fidelity 3D dance motions from monocular videos. To eliminate the physical artifacts prevalent in existing reconstruction methods, we introduce a Foot Restoration Diffusion Model (FRDM) guided by foot-contact and geometric constraints that enforce physical plausibility while preserving kinematic smoothness and expressiveness, resulting in a diverse, high-quality multimodal 3D dance dataset totaling 100.69 hours. (2) On model design, we propose Choreographic LLaMA (ChoreoLLaMA), a scalable LLaMA-based architecture. To enhance robustness under unfamiliar music conditions, we integrate a retrieval-augmented generation (RAG) module that injects reference dance as a prompt. Additionally, we design a slow/fast-cadence Mixture-of-Experts (MoE) module that enables ChoreoLLaMA to smoothly adapt motion rhythms across varying music tempos. Extensive experiments across diverse dance genres show that our approach surpasses existing methods in both qualitative and quantitative evaluations, marking a step toward scalable, real-world 3D dance generation. Code, models, and data will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。