用真人游戏数据构建大规模地理定位数据集,提升模型推理能力与准确性。
Geolocation with Real Human Gameplay Data: A Large-Scale Dataset and Human-Like Reasoning Framework
- 基于2500万条用户行为数据构建真实地理定位数据集
- 提出多步推理框架,使模型定位准确率提升25%
- 适合研究地理定位、视觉推理与人类行为建模的学者
地理定位任务需复杂推理,对导航、监控和文化保护至关重要。现有方法常产生粗略、不精确且不可解释的结果,主要因数据集规模小、质量差,且多为自动构建,导致数据噪声大、难度不一。为此,我们提出包含三部分的综合框架:GeoComp(大规模数据集)、GeoCoT(新型推理方法)和GeoEval(评估指标)。其中,GeoComp来自一个两年内覆盖740万用户的地理定位游戏平台,包含2500万条元数据和300万条地理标记位置,覆盖全球大部分地区,每处地点由数千至数万次人工标注。基于此数据集,我们提出地理链式思维(GeoCoT),一种模仿人类推理过程的多步视觉-空间融合方法,显著增强大视觉模型在地理定位中的推理能力。通过GeoEval评估,该方法将定位准确率最高提升25%,同时提高结果可解释性。
原文摘要 · Abstract (English)
Geolocation, the task of identifying an image's location, requires complex reasoning and is crucial for navigation, monitoring, and cultural preservation. However, current methods often produce coarse, imprecise, and non-interpretable localization. A major challenge lies in the quality and scale of existing geolocation datasets. These datasets are typically small-scale and automatically constructed, leading to noisy data and inconsistent task difficulty, with images that either reveal answers too easily or lack sufficient clues for reliable inference. To address these challenges, we introduce a comprehensive geolocation framework with three key components: GeoComp, a large-scale dataset; GeoCoT, a novel reasoning method; and GeoEval, an evaluation metric, collectively designed to address critical challenges and drive advancements in geolocation research. At the core of this framework is GeoComp (Geolocation Competition Dataset), a large-scale dataset collected from a geolocation game platform involving 740K users over two years. It comprises 25 million entries of metadata and 3 million geo-tagged locations spanning much of the globe, with each location annotated thousands to tens of thousands of times by human users. The dataset offers diverse difficulty levels for detailed analysis and highlights key gaps in current models. Building on this dataset, we propose Geographical Chain-of-Thought (GeoCoT), a novel multi-step reasoning framework designed to enhance the reasoning capabilities of Large Vision Models (LVMs) in geolocation tasks. GeoCoT improves performance by integrating contextual and spatial cues through a multi-step process that mimics human geolocation reasoning. Finally, using the GeoEval metric, we demonstrate that GeoCoT significantly boosts geolocation accuracy by up to 25% while enhancing interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。