用原子动作构建可解释的舞蹈生成框架,让舞步更连贯可控。
Music-to-Dance Generation via Atomic Movements

- 将舞蹈拆解为可解释的原子动作,作为生成基本单元
- 生成舞蹈结构更连贯,节奏对齐度提升显著
- 适合需要可编辑、可理解舞蹈生成的创作者使用
音乐驱动的舞蹈生成旨在创作与音乐节奏同步且语义一致的人体运动。现有神经方法虽在视觉真实感上表现优异,但通常将运动视为连续信号,忽略其组合特性,导致生成舞蹈结构松散、难以控制。本文提出一种结构感知框架,将编舞建模为一系列原子动作——具有语义意义的运动事件,是舞蹈的基本构成单元。首先,通过分割大规模舞蹈数据并聚类,构建原子动作词库;再利用大语言模型对聚类结果进行语义重标注与优化,获得可解释、可复用的原子动作集合。基于此,设计两阶段生成框架:第一阶段为原子动作规划,模型根据输入音乐预测动作类型、持续时间与时机,形成符号化舞蹈布局;第二阶段为补全阶段,采用过渡感知生成器,根据规划结构合成流畅且风格一致的运动。大量实验表明,相比现有基线,该方法生成的舞蹈在结构连贯性、节奏对齐和感知自然度上均有显著提升,同时通过显式结构表示实现了更强的可解释性与可控编辑能力。
原文摘要 · Abstract (English)
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its compositional nature, making generated dances structurally incoherent and difficult to control. In this work, we introduce a structure-aware framework that models choreography as a sequence of atomic movements-semantically interpretable motion events that serve as the building blocks of dance. To construct this atomic movement vocabulary, we first segment large-scale dance data and cluster them into atomic movement groups. We then employ a large language model to semantically relabel and refine the clusters, yielding a set of interpretable and reusable atomic movements. Based on these atomic movement annotations, we design a two-stage generation framework that mirrors the human choreography process. In the atomic movement planning stage, the model predicts the type, duration, and timing of atomic movements conditioned on the input music, forming a symbolic dance allocation. In the completion stage, a transition-aware generator synthesizes smooth and stylistically coherent motion conditioned on the planned structure. Extensive experiments demonstrate that our method produces dances with significantly improved structural coherence, rhythmic alignment, and perceptual naturalness compared to existing baselines, while providing enhanced interpretability and controllable editing through explicit structural representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。