解决遮挡下手物姿态估计难题,提升对未知物体的泛化能力。
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
- 用分层语义提示融合物体状态与手部先验,增强模型泛化能力。
- 在DexYCB和HO3Dv2上达到当前最优性能,尤其在严重遮挡下表现突出。
- 适合做通用手物交互建模、视觉理解与机器人操作的研究者参考。
单目RGB图像下的通用3D手物姿态估计仍面临巨大挑战,主要源于物体外观与交互模式的巨大差异,尤其是在严重遮挡条件下。本文提出GenHOI框架,通过引入具有遮挡感知能力的通用手物姿态估计方法。该框架将分层语义知识与手部先验相结合,以增强模型在复杂遮挡条件下的泛化能力。具体地,设计了一种分层语义提示,通过文本描述编码物体状态、手部构型及交互模式,使模型能够学习手物交互的抽象高层表征,从而实现对未见物体和新型交互的泛化,并补偿缺失或模糊的视觉线索。为实现鲁棒的遮挡推理,采用多模态掩码建模策略,覆盖RGB图像、预测点云和文本描述。此外,利用手部先验作为稳定的时空参考,提取隐式交互约束,支持在物体形状与交互模式显著变化时仍能可靠推断姿态。在具有挑战性的DexYCB与HO3Dv2基准测试中,实验表明本方法在手物姿态估计任务上达到了最先进的性能。
原文摘要 · Abstract (English)
Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under heavy occlusion. We propose GenHOI, a framework for generalized hand-object pose estimation with occlusion awareness. GenHOI integrates hierarchical semantic knowledge with hand priors to enhance model generalization under challenging occlusion conditions. Specifically, we introduce a hierarchical semantic prompt that encodes object states, hand configurations, and interaction patterns via textual descriptions. This enables the model to learn abstract high-level representations of hand-object interactions for generalization to unseen objects and novel interactions while compensating for missing or ambiguous visual cues. To enable robust occlusion reasoning, we adopt a multi-modal masked modeling strategy over RGB images, predicted point clouds, and textual descriptions. Moreover, we leverage hand priors as stable spatial references to extract implicit interaction constraints. This allows reliable pose inference even under significant variations in object shapes and interaction patterns. Extensive experiments on the challenging DexYCB and HO3Dv2 benchmarks demonstrate that our method achieves state-of-the-art performance in hand-object pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。