仅用多视角彩色图像实现纹理缺失物体的高精度6D姿态估计
DKPMV: Dense Keypoints Fusion from Multi-View RGB Frames for 6D Pose Estimation of Textureless Objects
- 三阶段渐进优化,融合多视角稠密关键点几何信息
- 在ROBI数据集上超越多数RGB方法,部分场景优于带深度的RGB-D方法
- 通过注意力聚合与对称性感知训练,提升对称物体估计精度
纹理缺失物体的6D姿态估计在工业机器人应用中至关重要,但因深度信息常丢失而难以实现。现有多视角方法或依赖深度数据,或未能充分挖掘多视角几何线索,性能受限。本文提出DKPMV,仅使用多视角彩色图像即可实现稠密关键点级融合。设计三阶段渐进姿态优化策略,充分利用稠密多视角关键点几何信息。为提升稠密关键点融合效果,增强关键点网络,引入注意力聚合与对称性感知训练,提高预测精度并解决对称物体的歧义问题。在ROBI数据集上的大量实验表明,DKPMV优于当前最先进的多视角RGB方法,且在多数情况下超过RGB-D方法。代码即将开源。
原文摘要 · Abstract (English)
6D pose estimation of textureless objects is valuable for industrial robotic applications, yet remains challenging due to the frequent loss of depth information. Current multi-view methods either rely on depth data or insufficiently exploit multi-view geometric cues, limiting their performance. In this paper, we propose DKPMV, a pipeline that achieves dense keypoint-level fusion using only multi-view RGB images as input. We design a three-stage progressive pose optimization strategy that leverages dense multi-view keypoint geometry information. To enable effective dense keypoint fusion, we enhance the keypoint network with attentional aggregation and symmetry-aware training, improving prediction accuracy and resolving ambiguities on symmetric objects. Extensive experiments on the ROBI dataset demonstrate that DKPMV outperforms state-of-the-art multi-view RGB approaches and even surpasses the RGB-D methods in the majority of cases. The code will be available soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。