解决视觉模型对地标依赖过强的问题,提升定位准确性
HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning

- 设计多证据推理框架,引导模型关注多样地理线索
- 在LandmarkBias-3K上超越现有开源模型,定位更可靠
- 适合需要鲁棒地理推理的研究者和应用开发者
视觉语言模型在图像地理定位方面取得进展,但易受地标偏差影响,忽略其他地理线索或形成虚假关联,导致定位不准。为此,我们提出两个量化指标——偏差强度(BI)与偏差危害性(BH),构建基准数据集LandmarkBias-3K。进一步提出HoloGeo框架,基于高质量的BF-30k数据集,包含结构化多证据无偏推理链,通过多维奖励机制,显式鼓励模型对多元视觉线索保持均衡关注,实现证据驱动的联合推理。大量实验表明,HoloGeo在IM2GPS3K和YFCC4k上保持优异性能,且在LandmarkBias-3K上显著优于现有开源VLM,验证其在鲁棒地理推理中的有效性。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization. To systematically investigate this issue, we first design two quantitative metrics, Bias Intensity (BI) and Bias Harmfulness (BH), to characterize the impact of landmarks exerted on model reasoning, and establish a comprehensive benchmark, LandmarkBias-3K. To mitigate landmark bias, we further propose an evidence-driven reasoning framework, HoloGeo, to improve the reliability of geo-localization. HoloGeo is supported by a high-quality dataset, BF-30k, annotated with structured multi-evidence bias-free reasoning chains. By incorporating multi-dimensional rewards, HoloGeo explicitly encourages balanced attention over diverse visual cues and achieves evidence-driven joint reasoning. Extensive experiments demonstrate that HoloGeo not only maintains excellent performance on IM2GPS3K and YFCC4k but also significantly outperforms existing open-source VLMs on LandmarkBias-3K, validating its effectiveness for robust geospatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。