让用户用文字逐步修改音乐驱动的舞蹈,生成更自然且可编辑的结果。
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- 采用先生成后编辑的框架,结合音乐与文本描述逐步优化动作。
- 在包含84.5K对数据的DanceRemix上实现优于现有模型的效果。
- 适合需要反复调整舞蹈动作的虚拟角色动画与编舞场景。
从音乐信号生成连贯多样的人类舞蹈在虚拟角色动画中已取得显著进展。然而,现有方法仅支持直接生成,无法满足用户在真实编舞场景中迭代修改动作的需求。此外,缺乏高质量的可编辑舞蹈数据集也制约了该问题的研究。为此,我们构建了DanceRemix——一个大规模多轮可编辑舞蹈数据集,包含超过2530万帧舞蹈动作和84500组配对数据。同时提出DanceEditor框架,实现与音乐同步、支持用户文本描述迭代编辑的舞蹈生成。该框架基于预测-编辑范式,融合多模态条件:初始阶段通过精准对齐的音乐直接建模舞蹈动作以提升生成质量;后续编辑阶段引入文本描述作为条件,通过专门设计的跨模态编辑模块(CEM)自适应融合初始预测、音乐与文本提示,作为时间运动线索引导合成序列。结果既保持音乐节奏和谐性,又精确匹配文本语义。大量实验表明,该方法在新构建的DanceRemix数据集上优于当前最先进模型。代码已公开于https://lzvsdy.github.io/DanceEditor/。
原文摘要 · Abstract (English)
Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multi-turn editable dance dataset comprising the prompt featuring over 25.3M dance frames and 84.5K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multi-modal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。