arXiv:2609.05761cs.CVcs.AI2026-09

提出新基准GeoContext,评估视觉定位中上下文依赖与位置验证的双重失败问题。

GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims

  • 构建分层上下文阶梯,模拟不同精度的位置提示进行定位测试。
  • 模型在远距离提示下误差显著上升,近处提示仅1个模型仍保持负准确率。
  • 定位验证任务中83.8%误通过以高置信度通过,暴露模型过度自信问题。

视觉地理定位基准通常不考虑用户提供的位置上下文。我们引入GeoContext,支持两个互补任务:给定真实但粗略的位置提示进行开放定位(GeoHint),以及二分类验证图像是否在声称地点150米范围内(GeoVerify)。GeoContext通过按距离和可参考性分层相邻参照点,固定图像而变化上下文。该基准覆盖30个城市中的109个站点,评估5个视觉语言模型,包含21,933次GeoHint响应和6,270次GeoVerify响应。评估揭示三个主要模式:首先,提示重复率在可参考性层级间仅差1.5个百分点,在距离带间差小于3个百分点,但定位误差随提示距离增加而持续上升;其次,行为强依赖无上下文性能:低无上下文准确率站点的中位误差与提示距离比约为1.00,高准确率站点则为0.24至0.69;经修正站点分组偏差后,仅一个模型在近提示下仍保持负准确率;第三,在GeoVerify中,无模型对150米外干扰项达到d' = 1;模型排名因敏感度与反应偏倚分离而改变,83.8%的误接受被以至少0.8置信度报告。我们发布基准、构建流程、审计决策及评分代码。

原文摘要 · Abstract (English)

Visual geolocation benchmarks typically ask a model where an image was captured without accounting for the location context that users often provide. We introduce GeoContext, a resource supporting two complementary tasks: GeoHint, open-ended localization given a true but coarse location hint, and GeoVerify, binary verification of whether an image was taken within 150 m of a claimed place. GeoContext constructs a context ladder by stratifying nearby reference points according to distance and referenceability, allowing the image to remain fixed while the supplied context varies. The benchmark covers 109 sites in 30 cities and evaluates five vision-language models using 21,933 GeoHint responses and 6,270 GeoVerify responses. Our evaluation reveals three main patterns. First, hint repetition varies by only 1.5 percentage points across referenceability tiers and by less than 3 points across distance bands, while the resulting localization error increases steadily with hint distance. Second, behavior depends strongly on no-context performance: at sites with low no-context accuracy, the median ratio between localization error and hint distance is approximately 1.00, whereas at higher-accuracy sites it ranges from 0.24 to 0.69. After correcting for bias introduced by the site grouping procedure, only one of the five models retains a negative accuracy estimate when given a nearby hint. Third, in GeoVerify, no model reaches d' = 1 for decoys immediately beyond the 150 m tolerance. Model rankings also change when sensitivity is separated from response bias, and 83.8% of false acceptances are reported with confidence of at least 0.8. We release the benchmark, construction pipeline, audit decisions, and scoring code.

视觉定位多模态基准测试上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。