arXiv:2503.06012cs.CV2025-03CVPR被引 9

用图编码的Transformer隐式建模人物交互,提升3D重建精度

End-to-End HOI Reconstruction Transformer with Graph-based Encoding

  • 通过自注意力机制隐式学习人体与物体的交互关系
  • 在InterCap数据集上人体和物体重建精度分别提升8.9%和8.6%
  • 无需复杂设计,端到端实现全局与局部结构的平衡

随着人-物交互(HOI)应用多样化及人体网格捕捉的成功,HOI重建受到广泛关注。现有主流方法通常显式建模人与物之间的交互,但这种做法导致3D网格重建(强调全局结构)与细粒度接触重建(关注局部细节)之间存在天然矛盾。为解决显式建模的局限性,本文提出端到端的基于图编码的HOI重建Transformer(HOI-TG)。该方法利用自注意力机制隐式学习人与物的交互,在Transformer架构中设计图残差块,聚合不同空间结构顶点间的拓扑关系,从而有效平衡全局与局部表征。无需额外复杂组件,HOI-TG在BEHAVE和InterCap数据集上均达到当前最优性能。尤其在具有挑战性的InterCap数据集上,人体和物体网格重建精度分别提升8.9%和8.6%。

原文摘要 · Abstract (English)

With the diversification of human-object interaction (HOI) applications and the success of capturing human meshes, HOI reconstruction has gained widespread attention. Existing mainstream HOI reconstruction methods often rely on explicitly modeling interactions between humans and objects. However, such a way leads to a natural conflict between 3D mesh reconstruction, which emphasizes global structure, and fine-grained contact reconstruction, which focuses on local details. To address the limitations of explicit modeling, we propose the End-to-End HOI Reconstruction Transformer with Graph-based Encoding (HOI-TG). It implicitly learns the interaction between humans and objects by leveraging self-attention mechanisms. Within the transformer architecture, we devise graph residual blocks to aggregate the topology among vertices of different spatial structures. This dual focus effectively balances global and local representations. Without bells and whistles, HOI-TG achieves state-of-the-art performance on BEHAVE and InterCap datasets. Particularly on the challenging InterCap dataset, our method improves the reconstruction results for human and object meshes by 8.9% and 8.6%, respectively.

HOI重建图神经网络3D生成Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。