用统一模型同时生成多人交互动作,提升质量与效率
Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer
- 将多人动作与互动关系统一压缩到潜空间,通过扩散过程生成
- 在文本条件驱动下,显著提升差异大动作的生成质量
- 适合需要高效生成复杂多人交互动作的动画场景
多人交互动作生成是角色动画中的关键但研究不足领域,面临建模人与人之间复杂互动关系及从同一文本条件生成差异巨大动作的挑战。现有方法通常采用独立模块处理个体动作,导致互动信息丢失且计算开销大。为此,我们提出一种新型统一方法,将多人动作及其互动关系建模于单一潜空间中。该方法通过变分自编码器(VAE)将交互动作压缩为统一潜变量,并在潜空间内进行扩散过程,由自然语言条件引导生成。实验表明,本方法在生成质量上优于现有方法,尤其在动作差异显著时表现更优,同时提升了生成效率并保持高质量输出。
原文摘要 · Abstract (English)
Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Current research often employs separate module branches for individual motions, leading to a loss of interaction information and increased computational demands. To address these challenges, we propose a novel, unified approach that models multi-person motions and their interactions within a single latent space. Our approach streamlines the process by treating interactive motions as an integrated data point, utilizing a Variational AutoEncoder (VAE) for compression into a unified latent space, and performing a diffusion process within this space, guided by the natural language conditions. Experimental results demonstrate our method's superiority over existing approaches in generation quality, performing text condition in particular when motions have significant asymmetry, and accelerating the generation efficiency while preserving high quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。