arXiv:2608.28386cs.CV2026-08

从单张照片重建人体与物体的3D交互,精确模拟手指抓握动作。

GraspHOI: Full-Body 3D Human-Object Reconstruction with Finger-Level Grasps from a Single In-the-Wild Image

论文配图:GraspHOI: Full-Body 3D Human-Object Reconstruction with Finger-Level Grasps from a Single In-the-Wild Image
图 1 · 摘自论文原文
  • 分离重建人体、手部和物体,通过深度与图像空间对齐
  • 显式优化手指与物体表面接触,避免穿透或悬空
  • 无需预设模型或类别词典,适用于任意物体

现有单目全身3D人-物交互(HOI)方法未将精细手指抓握优化与类别无关的物体重建结合。尽管人体-物体配置合理,其手指仍可能脱离或穿入物体。我们提出GraspHOI,首个从单张图像中重建全身3D HOI并显式优化手指关节以贴合重构物体的框架。该方法直接恢复物体几何形状,无需预定义网格或固定类别词典。身体、双手和物体分别重建,并通过基于深度的配准与图像空间对齐,在度量相机空间中统一。遮挡感知的掌面对应关系将物体定位在抓握手掌上,接触感知优化调整手臂与手指关节,确保表面接触且无过度穿透。在四个基准和六种基线中,GraspHOI提升了相对人-物位置、手部精度和接触合理性。完整代码将公开。

原文摘要 · Abstract (English)

Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic object reconstruction. Despite plausible body-object configurations, their fingers may float from or penetrate objects instead of forming a grasp. We present GraspHOI, the first framework that reconstructs a full-body 3D HOI from a single image while explicitly optimizing finger articulation against the reconstructed object. GraspHOI recovers object geometry directly, without predefined meshes or a fixed category vocabulary. It reconstructs the body, hands, and object separately, aligning them in metric camera space via depth-based registration and image-space alignment. Occlusion-aware palmar correspondences seat the object against the grasping hand, and contact-aware optimization refines arm and finger articulation to form surface contact without excessive penetration. Across four benchmarks and six baselines, GraspHOI improves relative human-object placement, hand accuracy, and contact plausibility. Full pipeline code will be released.

3D重建人机交互抓握识别单图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。