arXiv:2512.11811cs.CLcs.AI2025-12

用大模型增强图像定位,让市民拍的洪水照更快准确定位。

Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention

  • 用大模型识别关键地理区域,过滤无关干扰信息。
  • 在真实洪水图像上提升定位准确率最高达8%。
  • 无需重训练模型,适合应急响应和城市防灾场景。

众包社交媒体图像为城市洪涝提供了实时视觉证据,但常缺乏可靠的地理元数据,影响应急响应。现有视觉地点识别(VPR)模型因跨源域偏移和视觉失真难以有效定位。本文提出VPR-AttLLM,一种与模型无关的框架,通过注意力机制将大语言模型(LLMs)的语义推理与地理空间知识融入VPR流程,增强特征描述符。该方法利用LLM识别具地理意义的区域并抑制瞬时噪声,无需重新训练或新增数据即可提升检索效果。我们在旧金山和香港的数据集上评估,涵盖标准查询、模拟洪涝场景及真实社交媒体洪水图像。将VPR-AttLLM集成至先进模型(CosPlace、EigenPlaces、SALAD)后,召回率持续提升1-3%,在挑战性真实洪水图像上最高达8%。通过将城市感知原理嵌入注意力机制,该框架实现了人类式空间推理与现代VPR架构的融合。其即插即用设计与跨源鲁棒性,为快速定位众包危机影像提供可扩展解决方案,推动认知型城市韧性发展。

原文摘要 · Abstract (English)

Crowdsourced social media imagery provides real-time visual evidence of urban flooding but often lacks reliable geographic metadata for emergency response. Existing Visual Place Recognition (VPR) models struggle to geo-localize these images due to cross-source domain shifts and visual distortions. We present VPR-AttLLM, a model-agnostic framework integrating the semantic reasoning and geospatial knowledge of Large Language Models (LLMs) into VPR pipelines via attention-guided descriptor enhancement. VPR-AttLLM uses LLMs to isolate location-informative regions and suppress transient noise, improving retrieval without model retraining or new data. We evaluate this framework across San Francisco and Hong Kong using established queries, synthetic flooding scenarios, and real social media flood images. Integrating VPR-AttLLM with state-of-the-art models (CosPlace, EigenPlaces, SALAD) consistently improves recall, yielding 1-3% relative gains and up to 8% on challenging real flood imagery. By embedding urban perception principles into attention mechanisms, VPR-AttLLM bridges human-like spatial reasoning with modern VPR architectures. Its plug-and-play design and cross-source robustness offer a scalable solution for rapid geo-localization of crowdsourced crisis imagery, advancing cognitive urban resilience.

图像定位大模型应用应急响应城市防灾

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。