arXiv:2503.13303cs.CV2025-03CVPR被引 10

统一处理手部与手持物体的3D姿态估计,兼顾无物和有物场景。

UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation

论文配图:UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
图 1 · 摘自论文原文
  • 设计抓握感知特征融合模块与对象切换器,动态适应不同交互状态。
  • 通过合成去遮挡图像对提升模型在遮挡下的手姿估计鲁棒性。
  • 在三个基准上实现最优性能,适合多场景手部姿态应用。

从单目图像中估计手部及可能手持物体的3D姿态是一个长期挑战。现有方法通常专注于裸手或手物交互场景,难以灵活应对两者,且在非目标场景下性能下降。本文提出UniHOPE,一种统一的3D手物姿态估计方法,可灵活适应两种情形。技术上,设计了抓握感知特征融合模块,结合对象切换器,根据抓握状态动态控制手物姿态估计。为提升手部姿态估计在物体存在时的鲁棒性,构建真实去遮挡图像对,训练模型学习物体引起的遮挡模式,并提出多层次特征增强策略以学习不变于遮挡的特征。在三个常用基准上的大量实验表明,UniHOPE在手部独立与手物交互场景中均达到最先进水平。代码将发布于https://github.com/JoyboyWang/UniHOPE_Pytorch。

原文摘要 · Abstract (English)

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when applied to the other scenario. In this paper, we propose UniHOPE, a unified approach for general 3D hand-object pose estimation, flexibly adapting both scenarios. Technically, we design a grasp-aware feature fusion module to integrate hand-object features with an object switcher to dynamically control the hand-object pose estimation according to grasping status. Further, to uplift the robustness of hand pose estimation regardless of object presence, we generate realistic de-occluded image pairs to train the model to learn object-induced hand occlusions, and formulate multi-level feature enhancement techniques for learning occlusion-invariant features. Extensive experiments on three commonly-used benchmarks demonstrate UniHOPE's SOTA performance in addressing hand-only and hand-object scenarios. Code will be released on https://github.com/JoyboyWang/UniHOPE_Pytorch.

3D姿态估计手部建模视觉感知统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。