arXiv:2511.15874cs.CVcs.AI2025-11

针对遮挡下的6D位姿估计难题,提出四项改进方法提升精度与速度。

WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion

  • 动态非均匀采样聚焦可见区域,减少遮挡干扰。
  • 多假设推理保留多个候选位姿,避免单一路径失败。
  • 新增遮挡增强训练,适合机器人与AR场景应用。

准确的6D物体位姿估计对机器人、增强现实和场景理解至关重要。对于已见物体,通过逐物体微调可达到高精度,但泛化到未见物体仍是挑战。以往方法假设测试时可访问CAD模型,通常采用多阶段流程:检测分割物体、提出初始位姿、再进行优化。然而在遮挡情况下,早期阶段易出错,错误会沿序列处理传播,导致性能下降。为此,本文提出四项新改进:(i) 动态非均匀密集采样策略,聚焦可见区域,降低遮挡引发误差;(ii) 多假设推理机制,保留多个置信度排序的位姿候选,缓解单路径失效问题;(iii) 迭代优化逐步提升精度;(iv) 一系列面向遮挡的训练增强手段,增强鲁棒性与泛化能力。此外,提出基于可见性的加权评估指标,减少现有评测协议偏差。大量实验证明,本方法在ICBIN上精度提升超5%,BOP上提升超2%,且推理速度约快3倍。

原文摘要 · Abstract (English)

Accurate 6D object pose estimation is vital for robotics, augmented reality, and scene understanding. For seen objects, high accuracy is often attainable via per-object fine-tuning but generalizing to unseen objects remains a challenge. To address this problem, past arts assume access to CAD models at test time and typically follow a multi-stage pipeline to estimate poses: detect and segment the object, propose an initial pose, and then refine it. Under occlusion, however, the early-stage of such pipelines are prone to errors, which can propagate through the sequential processing, and consequently degrade the performance. To remedy this shortcoming, we propose four novel extensions to model-based 6D pose estimation methods: (i) a dynamic non-uniform dense sampling strategy that focuses computation on visible regions, reducing occlusion-induced errors; (ii) a multi-hypothesis inference mechanism that retains several confidence-ranked pose candidates, mitigating brittle single-path failures; (iii) iterative refinement to progressively improve pose accuracy; and (iv) series of occlusion-focused training augmentations that strengthen robustness and generalization. Furthermore, we propose a new weighted by visibility metric for evaluation under occlusion to minimize the bias in the existing protocols. Via extensive empirical evaluations, we show that our proposed approach achieves more than 5% improvement in accuracy on ICBIN and more than 2% on BOP dataset benchmarks, while achieving approximately 3 times faster inference.

6D位姿估计遮挡鲁棒模型基方法机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。