arXiv:2507.06486cs.CV2025-07中稿 · ICASSP 2024被引 15

用遮罩和3D-2D对应图提升单目图像中6D位姿估计精度

Mask6D: Masked Pose Priors For 6D Object Pose Estimation

  • 引入遮罩与3D-2D对应图作为额外模态进行预训练
  • 在复杂遮挡下相比现有方法提升位姿估计准确率
  • 适合需要高鲁棒性位姿估计的机器人抓取场景

在杂乱或遮挡条件下,仅用单目RGB图像实现鲁棒的6D物体位姿估计仍具挑战。原因在于当前位姿估计网络难以通过二维特征主干提取具有判别力的、与位姿相关的特征,尤其在目标被遮挡时信息不足。为此,我们提出一种名为Mask6D的新型位姿感知预训练策略。该方法将位姿相关的2D-3D对应图和可见掩码图作为额外模态,与RGB图像结合用于基于重建的模型预训练。其中,2D-3D对应图将变换后的3D物体模型映射到2D像素,反映目标在相机坐标系中的位姿信息;而集成的可见掩码图可有效引导模型忽略杂乱背景。此外,设计了面向物体的预训练损失函数,进一步帮助网络消除背景干扰。最后,采用常规位姿训练策略对预训练的位姿先验感知网络进行微调,实现可靠的位姿预测。大量实验验证,本方法优于以往端到端位姿估计方法。

原文摘要 · Abstract (English)

Robust 6D object pose estimation in cluttered or occluded conditions using monocular RGB images remains a challenging task. One reason is that current pose estimation networks struggle to extract discriminative, pose-aware features using 2D feature backbones, especially when the available RGB information is limited due to target occlusion in cluttered scenes. To mitigate this, we propose a novel pose estimation-specific pre-training strategy named Mask6D. Our approach incorporates pose-aware 2D-3D correspondence maps and visible mask maps as additional modal information, which is combined with RGB images for the reconstruction-based model pre-training. Essentially, this 2D-3D correspondence maps a transformed 3D object model to 2D pixels, reflecting the pose information of the target in camera coordinate system. Meanwhile, the integrated visible mask map can effectively guide our model to disregard cluttered background information. In addition, an object-focused pre-training loss function is designed to further facilitate our network to remove the background interference. Finally, we fine-tune our pre-trained pose prior-aware network via conventional pose training strategy to realize the reliable pose prediction. Extensive experiments verify that our method outperforms previous end-to-end pose estimation methods.

6D位姿估计单目视觉预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。