arXiv:2409.20502cs.LGcs.AI2024-09ICRA被引 4

用大模型引导扩散模型,生成逼真多人物协同动作。

COLLAGE: Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models

  • 结合大模型与分层潜变量自编码器,实现多层级动作建模。
  • 在CORE-4D和InterHuman数据集上生成动作更真实多样,性能领先。
  • 适合机器人、动画、视觉等领域复杂交互建模需求。

我们提出一种新框架COLLAGE,通过大语言模型(LLMs)与分层运动特定向量量化变分自编码器(VQ-VAEs)生成多人物协同物体交互动作。该模型利用LLM的知识与推理能力指导生成扩散模型,在潜空间中构建基于运动特征的分层表示,避免冗余概念并实现高效多分辨率建模。引入的扩散模型在潜空间中运行,并融合LLM生成的动作规划提示以引导去噪过程,实现更具控制性与多样性的提示驱动动作生成。在CORE-4D与InterHuman数据集上的实验表明,该方法能生成更真实、多样的协同人-物-人交互动作,优于现有先进方法。本工作为机器人、图形学与计算机视觉等领域建模复杂交互提供了新路径。

原文摘要 · Abstract (English)

We propose a novel framework COLLAGE for generating collaborative agent-object-agent interactions by leveraging large language models (LLMs) and hierarchical motion-specific vector-quantized variational autoencoders (VQ-VAEs). Our model addresses the lack of rich datasets in this domain by incorporating the knowledge and reasoning abilities of LLMs to guide a generative diffusion model. The hierarchical VQ-VAE architecture captures different motion-specific characteristics at multiple levels of abstraction, avoiding redundant concepts and enabling efficient multi-resolution representation. We introduce a diffusion model that operates in the latent space and incorporates LLM-generated motion planning cues to guide the denoising process, resulting in prompt-specific motion generation with greater control and diversity. Experimental results on the CORE-4D, and InterHuman datasets demonstrate the effectiveness of our approach in generating realistic and diverse collaborative human-object-human interactions, outperforming state-of-the-art methods. Our work opens up new possibilities for modeling complex interactions in various domains, such as robotics, graphics and computer vision.

动作生成扩散模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。