arXiv:2506.10816cs.CV2025-06被引 6

用掩码自编码器提升遮挡下手物姿态估计精度

Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders

  • 设计目标聚焦掩码策略,增强对遮挡区域的上下文感知
  • 融合SDF与点云,实现全局上下文与精细几何的互补
  • 在DexYCB和HO3Dv2上达到当前最优性能,适合工业应用

单目RGB图像下的手物姿态估计因手物交互中的严重遮挡仍具挑战性。现有方法对全局结构感知与推理利用不足,限制了其在遮挡场景下的表现。为此,我们提出基于掩码自编码器的遮挡感知手物姿态估计方法HOMAE。设计目标聚焦的掩码策略,在手物交互区域施加结构化遮挡,促使模型学习上下文感知特征并推断被遮挡结构。进一步融合解码器提取的多尺度特征,预测符号距离场(SDF),同时结合由SDF生成的显式点云,发挥两者互补优势。该融合策略通过SDF的全局上下文与点云的精确局部几何,更稳健地处理遮挡区域。在挑战性的DexYCB与HO3Dv2基准测试中,HOMAE实现了最先进的性能。代码与模型将公开。

原文摘要 · Abstract (English)

Hand-object pose estimation from monocular RGB images remains a significant challenge mainly due to the severe occlusions inherent in hand-object interactions. Existing methods do not sufficiently explore global structural perception and reasoning, which limits their effectiveness in handling occluded hand-object interactions. To address this challenge, we propose an occlusion-aware hand-object pose estimation method based on masked autoencoders, termed as HOMAE. Specifically, we propose a target-focused masking strategy that imposes structured occlusion on regions of hand-object interaction, encouraging the model to learn context-aware features and reason about the occluded structures. We further integrate multi-scale features extracted from the decoder to predict a signed distance field (SDF), capturing both global context and fine-grained geometry. To enhance geometric perception, we combine the implicit SDF with an explicit point cloud derived from the SDF, leveraging the complementary strengths of both representations. This fusion enables more robust handling of occluded regions by combining the global context from the SDF with the precise local geometry provided by the point cloud. Extensive experiments on challenging DexYCB and HO3Dv2 benchmarks demonstrate that HOMAE achieves state-of-the-art performance in hand-object pose estimation. We will release our code and model.

手物姿态掩码自编码器3D重建遮挡处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。