arXiv:2410.07164cs.CV2024-10被引 27

零样本生成真人与物体互动的4D动画,直接从文字描述出发。

AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation

  • 用大模型识别文本中的接触部位,精准定位人体与物体空间关系。
  • 通过对应关系优化运动场,避免穿模,生成自然连贯的动作。
  • 无需真实数据训练,适合快速创建日常场景中的人物互动动画。

扩散模型在4D全身人-物交互(HOI)生成与动画方面取得显著进展,但现有方法多基于SMPL人体模型,受限于真实大规模交互数据稀缺,难以生成日常生活中的自然人-物交互场景。本文提出AvatarGO,一种基于预训练扩散模型的零样本框架,可直接从文本输入生成可动画化的4D HOI场景。针对“何处交互”挑战,引入LLM引导的接触重定向机制,利用Lang-SAM从文本提示中识别接触身体部位,确保人体与物体空间关系精确表达;针对“如何交互”挑战,提出对应感知运动优化,基于SMPL-X的线性混合皮肤函数构建人体与物体的运动场,实现协同动作生成并显著提升对穿模问题的鲁棒性。大量实验表明,AvatarGO在多种人-物组合和姿态下均优于现有方法,在生成质量与动画稳定性上表现优异。作为首个实现带物交互4D角色合成的工作,AvatarGO有望推动以人为中心的4D内容创作新范式。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation, which is limited by the scarcity of realistic large-scale interaction data. This constraint affects their ability to create everyday HOI scenes. This paper addresses this challenge using a zero-shot approach with a pre-trained diffusion model. Despite this potential, achieving our goals is difficult due to the diffusion model's lack of understanding of ''where'' and ''how'' objects interact with the human body. To tackle these issues, we introduce AvatarGO, a novel framework designed to generate animatable 4D HOI scenes directly from textual inputs. Specifically, 1) for the ''where'' challenge, we propose LLM-guided contact retargeting, which employs Lang-SAM to identify the contact body part from text prompts, ensuring precise representation of human-object spatial relations. 2) For the ''how'' challenge, we introduce correspondence-aware motion optimization that constructs motion fields for both human and object models using the linear blend skinning function from SMPL-X. Our framework not only generates coherent compositional motions, but also exhibits greater robustness in handling penetration issues. Extensive experiments with existing methods validate AvatarGO's superior generation and animation capabilities on a variety of human-object pairs and diverse poses. As the first attempt to synthesize 4D avatars with object interactions, we hope AvatarGO could open new doors for human-centric 4D content creation.

4D生成人物交互零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。