用游戏数据训练视觉语言模型,让图像定位更准且少样本
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
- 用地理游戏生成专家推理数据,指导模型理解跨模态上下文
- 仅用不到1000张图训练,定位误差降低14%
- 适合研究少样本定位与多模态推理的学者
图像地理定位需综合视觉、地理与文化背景进行复杂推理。尽管先前的视觉语言模型(VLMs)在此任务上表现最佳,但高质量数据集与分析型模型仍十分稀缺。我们首先基于流行地理游戏GeoGuessr构建了高质量数据集NaviClues,提供语言引导的专家推理示例。基于此,提出Navig框架,融合全局与细粒度图像信息,通过语言推理将平均定位误差降低14%,且仅需少于1000个训练样本。相关数据集与代码已开源:https://github.com/SparrowZheyuan18/Navig/
原文摘要 · Abstract (English)
Image geo-localization is the task of predicting the specific location of an image and requires complex reasoning across visual, geographical, and cultural contexts. While prior Vision Language Models (VLMs) have the best accuracy at this task, there is a dearth of high-quality datasets and models for analytical reasoning. We first create NaviClues, a high-quality dataset derived from GeoGuessr, a popular geography game, to supply examples of expert reasoning from language. Using this dataset, we present Navig, a comprehensive image geo-localization framework integrating global and fine-grained image information. By reasoning with language, Navig reduces the average distance error by 14% compared to previous state-of-the-art models while requiring fewer than 1000 training samples. Our dataset and code are available at https://github.com/SparrowZheyuan18/Navig/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。