用生成式掩码先验实现高保真可编辑的舞蹈动作合成
Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
- 基于文本到动作的掩码模型构建分布先验,融合音乐、风格和姿态多信号
- 在Dance-4D数据集上达到新最高分,运动真实感与音乐同步性显著提升
- 支持动作补全和身体部位修改,适合动画师与编舞创作者使用
近期舞蹈生成进展实现了3D舞蹈动作的自动合成,但现有方法仍难以同时保证高真实感、精准的舞乐同步、多样化的动作表达和物理合理性。为此,我们提出一种新方法,利用生成式掩码文本到动作模型作为分布先验,学习从音乐、风格和姿态等多种引导信号到高质量舞蹈序列的概率映射。框架还支持语义动作编辑,如动作补全和身体部位修改。具体地,我们设计了一个多塔掩码动作模型,包含文本条件的掩码动作主干以及并行的音乐引导塔和姿态引导塔。模型通过同步渐进式掩码训练进行训练,有效注入预训练的文本到动作先验,同时使各引导分支独立优化,避免梯度干扰。推理时引入无分类器逻辑引导和姿态引导的标记优化,增强音乐、风格和姿态信号的影响。大量实验表明,该方法在舞蹈生成方面达到新最优性能,显著提升质量和可编辑性。项目页面见https://foram-s1.github.io/DanceMosaic/
原文摘要 · Abstract (English)
Recent advances in dance generation have enabled the automatic synthesis of 3D dance motions. However, existing methods still face significant challenges in simultaneously achieving high realism, precise dance-music synchronization, diverse motion expression, and physical plausibility. To address these limitations, we propose a novel approach that leverages a generative masked text-to-motion model as a distribution prior to learn a probabilistic mapping from diverse guidance signals, including music, genre, and pose, into high-quality dance motion sequences. Our framework also supports semantic motion editing, such as motion inpainting and body part modification. Specifically, we introduce a multi-tower masked motion model that integrates a text-conditioned masked motion backbone with two parallel, modality-specific branches: a music-guidance tower and a pose-guidance tower. The model is trained using synchronized and progressive masked training, which allows effective infusion of the pretrained text-to-motion prior into the dance synthesis process while enabling each guidance branch to optimize independently through its own loss function, mitigating gradient interference. During inference, we introduce classifier-free logits guidance and pose-guided token optimization to strengthen the influence of music, genre, and pose signals. Extensive experiments demonstrate that our method sets a new state of the art in dance generation, significantly advancing both the quality and editability over existing approaches. Project Page available at https://foram-s1.github.io/DanceMosaic/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。