融合视觉与深度信息,提升物体实例重识别精度,实现跨视角精准定位。
Towards Global Localization using Multi-Modal Object-Instance Re-Identification
- 设计双路径Transformer架构,融合RGB与深度多模态信息。
- 在复杂场景下实现75.18 mAP的重识别准确率,定位成功率达83%。
- 适用于机器人长期感知与自主探索,数据集与代码开源。
重识别(ReID)是计算机视觉中的关键挑战,以往主要聚焦于行人与车辆。然而,对物体实例重识别的鲁棒性研究仍不充分,而这一能力对自主探索、长期感知和场景理解具有重要意义。本文提出一种新型双路径物体实例重识别变换器架构,融合多模态RGB与深度信息。利用深度数据,我们在杂乱或光照变化场景中显著提升了重识别性能。此外,构建基于重识别的定位框架,实现不同视角下的相机定位与姿态识别。通过自建两个RGB-D数据集及TUM RGB-D公开数据集中的多个序列验证方法有效性。实验表明,该方法在物体实例重识别(mAP达75.18)和定位精度(TUM-RGBD上成功率达83%)方面均有显著提升,凸显了物体重识别在推进机器人感知中的核心作用。模型、框架与数据集均已公开。
原文摘要 · Abstract (English)
Re-identification (ReID) is a critical challenge in computer vision, predominantly studied in the context of pedestrians and vehicles. However, robust object-instance ReID, which has significant implications for tasks such as autonomous exploration, long-term perception, and scene understanding, remains underexplored. In this work, we address this gap by proposing a novel dual-path object-instance re-identification transformer architecture that integrates multimodal RGB and depth information. By leveraging depth data, we demonstrate improvements in ReID across scenes that are cluttered or have varying illumination conditions. Additionally, we develop a ReID-based localization framework that enables accurate camera localization and pose identification across different viewpoints. We validate our methods using two custom-built RGB-D datasets, as well as multiple sequences from the open-source TUM RGB-D datasets. Our approach demonstrates significant improvements in both object instance ReID (mAP of 75.18) and localization accuracy (success rate of 83% on TUM-RGBD), highlighting the essential role of object ReID in advancing robotic perception. Our models, frameworks, and datasets have been made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。