arXiv:2412.14706cs.CV2024-12CVPR被引 33

用能量模型融合多语义,生成更连贯的复杂动作序列。

EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space

  • 将扩散模型视为潜空间中的能量模型,通过组合多个扩散模型生成动作。
  • 引入语义感知能量模型,实现文本嵌入的自适应梯度下降与语义融合。
  • 适合需要多概念组合生成的动画、游戏动作设计场景。

扩散模型,尤其是潜空间扩散模型,在文本驱动的人体动作生成中表现卓越。然而,潜空间扩散模型在将多个语义概念整合为单一连贯动作序列方面仍面临挑战。为此,我们提出EnergyMoGen,包含两种能量模型:(1) 将扩散模型视为潜空间感知的能量模型,通过组合多个扩散模型生成动作;(2) 引入基于交叉注意力的语义感知能量模型,实现语义组合与文本嵌入的自适应梯度下降。为解决两种模型谱之间的语义不一致与动作畸变问题,我们设计了协同能量融合机制。该机制使运动潜空间扩散模型能结合对应文本描述的多个能量项,合成高质量、复杂的动作序列。实验表明,我们的方法在多种任务上优于现有最先进模型,包括文本到动作生成、组合动作生成和多概念动作生成。此外,我们还证明该方法可用于扩展动作数据集并提升文本到动作任务性能。

原文摘要 · Abstract (English)

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic concepts into a single, coherent motion sequence. To address this issue, we propose EnergyMoGen, which includes two spectrums of Energy-Based Models: (1) We interpret the diffusion model as a latent-aware energy-based model that generates motions by composing a set of diffusion models in latent space; (2) We introduce a semantic-aware energy model based on cross-attention, which enables semantic composition and adaptive gradient descent for text embeddings. To overcome the challenges of semantic inconsistency and motion distortion across these two spectrums, we introduce Synergistic Energy Fusion. This design allows the motion latent diffusion model to synthesize high-quality, complex motions by combining multiple energy terms corresponding to textual descriptions. Experiments show that our approach outperforms existing state-of-the-art models on various motion generation tasks, including text-to-motion generation, compositional motion generation, and multi-concept motion generation. Additionally, we demonstrate that our method can be used to extend motion datasets and improve the text-to-motion task.

动作生成扩散模型语义融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。