arXiv:2506.01277cs.AI2025-06被引 13

用少量高质量数据微调大模型,实现高效精准的图像地理定位。

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

  • 仅用2700对图像-位置数据,通过监督微调提升定位性能。
  • 在Im2GPS-3k、YFCC-4k和新提出的MR40k基准上表现优异。
  • 适合关注高效地理定位与小样本训练的研究者使用。

准确确定单张图像拍摄的地理位置(视觉地理定位)仍具挑战性,因地球范围广阔且遥远地点外观相似。我们提出GeoLocSFT框架,展示如何通过少量高质量数据对大型多模态基础模型(Gemma 3)进行针对性监督微调(SFT),即可获得极具竞争力的定位性能。该方法仅使用我们地理多样化的MR600k数据集中精心挑选的2700个图像-GPS配对进行训练。尽管数据量有限,其微调策略显著优于基线模型,在标准基准如Im2GPS-3k、YFCC-4k以及新提出的、针对稀疏区域设计的挑战性MR40k基准上均取得稳健结果。我们还探索了多候选推理与聚合策略,但发现核心提升已体现在SFT阶段。研究凸显高质量监督与高效SFT在行星尺度图像定位中的潜力,尤其相较于以往需海量数据库或复杂流程的方法。为促进后续研究,我们公开发布MR40k基准数据集。

原文摘要 · Abstract (English)

Accurately determining the geographic location where a single image was taken, visual geolocation, remains a formidable challenge due to the planet's vastness and the deceptive similarity among distant locations. We introduce GeoLocSFT, a framework that demonstrates how targeted supervised fine-tuning (SFT) of a large multimodal foundation model (Gemma 3) using a small, high-quality dataset can yield highly competitive geolocation performance. GeoLocSFT is trained with only 2700 carefully selected image-GPS pairs from our geographically diverse MR600k dataset. Despite this limited data, our SFT-centric approach substantially improves over baseline models and achieves robust results on standard benchmarks such as Im2GPS-3k and YFCC-4k, as well as on our newly proposed and challenging MR40k benchmark, aimed specifically at sparsely populated regions. Further, we explore multi-candidate inference and aggregation strategies but find that the core gains are already realized at the SFT stage. Our findings highlight the power of high-quality supervision and efficient SFT for planet-scale image geolocation, especially when compared to prior methods that require massive databases or complex pipelines. To foster further research, we publicly release the MR40k benchmark dataset.

视觉定位多模态模型微调地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。