arXiv:2512.22324cs.CV2025-12

用能量扩散模型分解复杂动作,自动发现可复用的运动基元。

DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models

  • 基于能量扩散模型,从完整动作中无监督分解出语义子成分。
  • 三种训练变体使模型在无真值标注下仍能准确拆解动作概念。
  • 分解后的运动基元可灵活重组合,生成新颖多样动作,适合动画设计者使用。

人类动作具有组合性:复杂行为可由简单基元组合而成。现有方法多聚焦正向建模,如从文本生成完整动作或由多个动作概念拼接而成。本文提出反向视角:将整体动作分解为语义明确的子成分。我们设计DeMoGen,一种基于能量型扩散模型的组合式训练范式,直接建模多个动作概念的联合分布,使模型无需依赖单个概念的真实动作即可发现其结构。该范式包含三种训练策略:DeMoGen-Exp 显式使用分解后的文本提示;DeMoGen-OSS 实现正交自监督分解;DeMoGen-SC 确保原始与分解后文本嵌入的语义一致性。这些方法使模型能有效解耦出可复用的运动基元。我们还构建了一个文本分解数据集,支持组合式训练,扩展了文本到动作生成与动作组合的能力。实验表明,分解出的动作概念可灵活重组,生成超越训练分布的新颖动作。

原文摘要 · Abstract (English)

Human motions are compositional: complex behaviors can be described as combinations of simpler primitives. However, existing approaches primarily focus on forward modeling, e.g., learning holistic mappings from text to motion or composing a complex motion from a set of motion concepts. In this paper, we consider the inverse perspective: decomposing a holistic motion into semantically meaningful sub-components. We propose DeMoGen, a compositional training paradigm for decompositional learning that employs an energy-based diffusion model. This energy formulation directly captures the composed distribution of multiple motion concepts, enabling the model to discover them without relying on ground-truth motions for individual concepts. Within this paradigm, we introduce three training variants to encourage a decompositional understanding of motion: 1. DeMoGen-Exp explicitly trains on decomposed text prompts; 2. DeMoGen-OSS performs orthogonal self-supervised decomposition; 3. DeMoGen-SC enforces semantic consistency between original and decomposed text embeddings. These variants enable our approach to disentangle reusable motion primitives from complex motion sequences. We also demonstrate that the decomposed motion concepts can be flexibly recombined to generate diverse and novel motions, generalizing beyond the training distribution. Additionally, we construct a text-decomposed dataset to support compositional training, serving as an extended resource to facilitate text-to-motion generation and motion composition.

动作生成扩散模型分解学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。