arXiv:2410.14508cs.CVcs.AI2024-10被引 11

让文字描述更精准地生成自然动作,提升真实感与语义对齐。

LEAD: Latent Realignment for Human Motion Diffusion

  • 通过潜空间重校准,构建语义结构化的动作表征空间。
  • 在HumanML3D和KIT-ML数据集上达到顶尖的逼真度与一致性表现。
  • 适合需要高语义对齐动作生成或从少量样本学习新动作的研究者。

本文旨在从自然语言生成逼真的人体动作。现有方法常面临模型表达能力与文本-动作对齐之间的权衡:部分方法虽对齐了文本与动作潜空间,但牺牲了表达多样性;另一些依赖扩散模型生成惊艳动作,但其潜空间缺乏语义意义,可能影响真实感、多样性与实用性。为此,我们结合潜空间扩散与重校准机制,构建了一个新型的语义结构化空间,可编码语言语义。基于此,我们提出文本动作逆向任务,仅需少数样例即可捕捉新动作概念。在HumanML3D和KIT-ML上的实验表明,LEAD在逼真度、多样性和文本-动作一致性方面达到当前最优水平。定性分析与用户研究显示,合成动作更清晰、更类人,且更贴合文本描述。在文本动作逆向任务中,相比传统VAEs,本方法对分布外特征的建模能力显著提升。

原文摘要 · Abstract (English)

Our goal is to generate realistic human motion from natural language. Modern methods often face a trade-off between model expressiveness and text-to-motion alignment. Some align text and motion latent spaces but sacrifice expressiveness; others rely on diffusion models producing impressive motions, but lacking semantic meaning in their latent space. This may compromise realism, diversity, and applicability. Here, we address this by combining latent diffusion with a realignment mechanism, producing a novel, semantically structured space that encodes the semantics of language. Leveraging this capability, we introduce the task of textual motion inversion to capture novel motion concepts from a few examples. For motion synthesis, we evaluate LEAD on HumanML3D and KIT-ML and show comparable performance to the state-of-the-art in terms of realism, diversity, and text-motion consistency. Our qualitative analysis and user study reveal that our synthesized motions are sharper, more human-like and comply better with the text compared to modern methods. For motion textual inversion, our method demonstrates improved capacity in capturing out-of-distribution characteristics in comparison to traditional VAEs.

动作生成扩散模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。