arXiv:2505.06575cs.CV2025-05被引 4

从2D图像精准定位3D人体与场景接触点,提升交互理解能力。

GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images

  • 用点云编码解码结构融合3D人体几何与2D图像语义
  • 在多个数据集上达到当前最优接触点预测精度
  • 适用于不同体型人体,泛化能力强,适合虚拟现实应用

估计人体与场景的几何级接触旨在将特定接触点定位到3D人体几何体上,提供空间先验并连接人体与场景的交互,支持行为分析、具身AI和AR/VR等应用。现有方法多依赖参数化人体模型(如SMPL),通过固定的SMPL顶点序列建立图像与接触区域的对应关系,实质上完成了从图像特征到有序序列的映射。但该方法忽略几何多样性,限制了在不同人体形态下的泛化能力。本文提出GRACE(Geometry-level Reasoning for 3D Human-scene Contact Estimation),一种新的3D人体接触估计范式。GRACE采用点云编码器-解码器架构,结合层次化特征提取与融合模块,有效整合3D人体几何结构与来自图像的2D交互语义。在视觉线索引导下,建立几何特征到3D人体网格顶点空间的隐式映射,从而实现接触区域的精确建模。该设计确保高预测精度,并赋予框架对多样化人体几何的强大泛化能力。大量实验在多个基准数据集上表明,GRACE在接触估计任务中达到最先进性能,额外结果进一步验证其对非结构化人体点云的鲁棒泛化能力。

原文摘要 · Abstract (English)

Estimating the geometry level of human-scene contact aims to ground specific contact surface points at 3D human geometries, which provides a spatial prior and bridges the interaction between human and scene, supporting applications such as human behavior analysis, embodied AI, and AR/VR. To complete the task, existing approaches predominantly rely on parametric human models (e.g., SMPL), which establish correspondences between images and contact regions through fixed SMPL vertex sequences. This actually completes the mapping from image features to an ordered sequence. However, this approach lacks consideration of geometry, limiting its generalizability in distinct human geometries. In this paper, we introduce GRACE (Geometry-level Reasoning for 3D Human-scene Contact Estimation), a new paradigm for 3D human contact estimation. GRACE incorporates a point cloud encoder-decoder architecture along with a hierarchical feature extraction and fusion module, enabling the effective integration of 3D human geometric structures with 2D interaction semantics derived from images. Guided by visual cues, GRACE establishes an implicit mapping from geometric features to the vertex space of the 3D human mesh, thereby achieving accurate modeling of contact regions. This design ensures high prediction accuracy and endows the framework with strong generalization capability across diverse human geometries. Extensive experiments on multiple benchmark datasets demonstrate that GRACE achieves state-of-the-art performance in contact estimation, with additional results further validating its robust generalization to unstructured human point clouds.

3D人体建模接触估计几何推理点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。