用视觉语言模型提升物体地图定位的匹配精度和鲁棒性
CLIP-Clique: Graph-based Correspondence Matching Augmented by Vision Language Models for Object-based Global Localization
- 用VLM增强地标描述符,减少误分类和遮挡影响
- 通过图论方法确定性地识别内点,避免RANSAC随机性
- 结合相似度与观测完整性加权最小二乘,提升姿态估计精度
本文提出一种基于语义物体地标的全局定位方法。现有方法依赖周围物体分布生成的地标描述符进行语义图匹配,但易受误分类和部分观测影响。同时,多数方法使用随机性强、对高外点率敏感的RANSAC进行内点提取。为此,本文引入视觉语言模型(VLM)增强对应匹配:利用独立于周围物体的VLM嵌入提升地标可区分性;采用图论方法确定性估计内点;并结合对应相似度与观测完整性,通过加权最小二乘计算位姿,提升鲁棒性。在ScanNet和TUM数据集上的实验验证了匹配与位姿估计精度的提升。
原文摘要 · Abstract (English)
This letter proposes a method of global localization on a map with semantic object landmarks. One of the most promising approaches for localization on object maps is to use semantic graph matching using landmark descriptors calculated from the distribution of surrounding objects. These descriptors are vulnerable to misclassification and partial observations. Moreover, many existing methods rely on inlier extraction using RANSAC, which is stochastic and sensitive to a high outlier rate. To address the former issue, we augment the correspondence matching using Vision Language Models (VLMs). Landmark discriminability is improved by VLM embeddings, which are independent of surrounding objects. In addition, inliers are estimated deterministically using a graph-theoretic approach. We also incorporate pose calculation using the weighted least squares considering correspondence similarity and observation completeness to improve the robustness. We confirmed improvements in matching and pose estimation accuracy through experiments on ScanNet and TUM datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。