arXiv:2605.04554cs.CV2026-05

显式建模人与环境互动,提升多人姿态重建精度。

InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery

论文配图:InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery
图 1 · 摘自论文原文
  • 引入交互检测器,显式注入人-物、人-人交互语义。
  • 在CMU Panoptic和Hi4D上分别降低9.9%和8.2%的MPJPE。
  • 轻量模块可无缝集成至现有姿态重建框架,适合复杂场景应用。

人类持续与环境互动。现有基于DETR框架的端到端多人网格重建方法通过跨所有人体查询的自注意力捕捉人际关联,但仅隐式建模交互,缺乏对人与物体及他人互动的显式推理。本文提出InterMesh,一种简单而有效的框架,将人-环境互动信息显式融入网格重建流程。通过人-物交互检测器,InterMesh丰富查询表征中的结构化交互语义,提升姿态与形状估计精度。设计轻量级模块:上下文交互编码器与交互引导精修器,以极低开销整合至现有HMR架构。在3DPW、MuPoTS、CMU Panoptic、Hi4D和CHI3D数据集上广泛验证,显著优于当前最优方法。特别地,在CMU Panoptic上MPJPE降低9.9%,在Hi4D上降低8.2%,凸显其在复杂人-物与人际互动场景下的有效性。代码与模型已公开于https://github.com/Kelly510/InterMesh。

原文摘要 · Abstract (English)

Humans constantly interact with their surroundings. Existing end-to-end multi-person human mesh recovery methods, typically based on the DETR framework, capture inter-human relationships through self-attention across all human queries. However, these approaches model interactions only implicitly and lack explicit reasoning about how humans interact with objects and with each other. In this paper, we propose InterMesh, a simple yet effective framework that explicitly incorporates human-environment interaction information into human mesh recovery pipeline. By leveraging a human-object interaction detector, InterMesh enriches query representations with structured interaction semantics, enabling more accurate pose and shape estimation. We design lightweight modules, Contextual Interaction Encoder and Interaction-Guided Refiner, to integrate these features into existing HMR architectures with minimal overhead. We validate our approach through extensive experiments on 3DPW, MuPoTS, CMU Panoptic, Hi4D, and CHI3D datasets, demonstrating remarkable improvements over state-of-the-art methods. Notably, InterMesh reduces MPJPE by 9.9% on CMU Panoptic and 8.2% on Hi4D, highlighting its effectiveness in scenarios with complex human-object and inter-human interactions. Code and models are released at https://github.com/Kelly510/InterMesh.

姿态估计人机交互网格重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。