arXiv:2503.13130cs.CV2025-03CVPR被引 20

通过关节与运动链建模,生成更真实的文本驱动人物交互动作

ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation

  • 用关节图和时空图卷积网络显式建模关节间交互
  • 引入运动学模块,提升动作的生物力学合理性
  • 适合需要高真实感动作生成的研究者与开发者

我们提出 ChainHOI,一种新型的文本驱动人-物交互(HOI)生成方法,显式地在关节层级和运动链层级建模交互。与现有方法将全身姿态作为隐变量建模不同,我们认为显式建模关节级交互更自然有效,能直接捕捉关节间的几何与语义关系,而非在潜在姿态空间中建模。为此,ChainHOI引入新型关节图以捕获与物体的潜在交互,并采用生成式时空图卷积网络在关节层面显式建模交互。此外,我们提出基于运动学的交互模块,在运动链层级显式建模交互,确保动作更具真实性和生物力学一致性。在两个公开数据集上的评估表明,ChainHOI 显著优于以往方法,生成的动作更真实、语义一致。代码已开源。

原文摘要 · Abstract (English)

We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly modeling joint-level interactions is more natural and effective for generating realistic HOIs, as it directly captures the geometric and semantic relationships between joints, rather than modeling interactions in the latent pose space. To this end, ChainHOI introduces a novel joint graph to capture potential interactions with objects, and a Generative Spatiotemporal Graph Convolution Network to explicitly model interactions at the joint level. Furthermore, we propose a Kinematics-based Interaction Module that explicitly models interactions at the kinetic chain level, ensuring more realistic and biomechanically coherent motions. Evaluations on two public datasets demonstrate that ChainHOI significantly outperforms previous methods, generating more realistic, and semantically consistent HOIs. Code is available \href{https://github.com/qinghuannn/ChainHOI}{here}.

动作生成人机交互图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。