测试顶级视觉语言模型在无训练条件下的全球定位能力,发现其粗略定位强但精细定位差。
Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?
- 用文本提示直接调用黑盒视觉语言模型进行零样本定位
- 在真实约束下,粗粒度定位准确率高,细粒度定位明显下降
- 提出一致性指标评估生成模型的稳定性,适合开放世界机器人导航研究者
视觉语言模型(VLMs)的发展为仅基于视觉数据的图像地理定位提供了新可能,尤其适用于机器人在未知位置时的全局重定位。现有工作多将VLM作为嵌入提取器使用,但最先进的VLM常以黑盒API形式提供,无法访问训练数据、特征或梯度,且预测次数受限。本文首次系统性研究了先进生成式VLM在黑盒条件下作为独立零样本地理定位系统在全星球尺度上的潜力,考察三种场景:固定文本提示、语义等价文本提示、语义等价查询图像。除标准精度外,引入模型一致性作为衡量生成模型自回归与概率特性的重要指标。结果表明,尽管VLM具备较强的粗粒度定位与导航先验,但在真实变化下细粒度定位性能显著下降,揭示其在鲁棒开放世界机器人导航中部署的可靠性挑战。
原文摘要 · Abstract (English)
The advances in Vision-Language models (VLMs) offer exciting opportunities for robotic applications involving image geo-localization - the problem of identifying the geo-coordinates of a place based on visual data only. In robotics, such capabilities are particularly relevant to the global re-localization stage of the kidnapped robot problem, where a robot must recover its pose without prior knowledge of its location. Recent work has focused on using a VLM as embedding extractor for geo-localization. However, the most sophisticated VLMs may only be available as black boxes that are accessible through an API, and come with a number of limitations: there is no access to training data, model features and gradients; retraining is not possible; and the number of predictions may be limited by the API. The potential of state-of-the-art VLMs as a stand-alone, zero-shot geo-localization systems at planet scale using a single text-based prompt is largely unexplored. To bridge this gap, this paper undertakes the first systematic study, to the best of our knowledge, to investigate state-of-the-art generative VLMs as stand-alone, zero-shot geo-localization systems in a black-box setting with realistic constraints. We consider three main scenarios for this thorough investigation: a) fixed text-based prompt; b) semantically-equivalent text-based prompts; and c) semantically-equivalent query images. Beyond standard accuracy, we introduce model consistency as a metric to account for the auto-regressive and probabilistic nature of generative VLMs. Our findings reveal that while VLMs demonstrate strong coarse-level localization and navigation priors, fine-grained localization degrades significantly under realistic variations, highlighting reliability challenges for deploying generative VLMs in robust, open-world robotic navigation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。