arXiv:2607.24856cs.CVcs.AI2026-07

用多模态大模型+跨视角匹配,精准定位社交媒体中的灾情地名。

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

论文配图:DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
图 1 · 摘自论文原文
  • 结合多模态大模型生成候选位置,再通过多源影像交叉验证。
  • 在哈维飓风数据集上,50米内定位准确率达47.01%,误差中位数仅0.68公里。
  • 尤其擅长处理模糊地名,显著减少位置歧义和误差,适合应急响应场景。

社交媒体图像(SMI)能提供及时且精细的灾情地面视角,对态势感知和应急响应具有重要价值。与卫星或航拍图像不同,SMI可实时捕捉灾害影响和地面状况。然而,其地理引用常模糊或歧义,导致精确定位困难。为此,本文提出DisasterTD框架,融合多模态大语言模型(MLLM)语义推理与跨视角地理定位。首先,MLLM从嘈杂文本中提取地名并生成候选地理位置;随后,通过社交媒体图像(SMI)、遥感影像(RSI)及可选街景图像(SVI)间的跨视图匹配,验证并优化候选结果。我们在哈维飓风数据集上评估该方法,构建了包含采集的RSI和SVI的跨视图基准,按地名清晰度分为四类,实现细粒度性能分析。结果显示,DisasterTD持续优于仅使用MLLM或仅使用跨视图的基线,达到1000米内71.62%、500米内62.36%、250米内57.99%、100米内52.09%、50米内47.01%的定位准确率,平均误差降至11.33公里,中位误差为0.68公里。在模糊地名场景下,语义推理结合跨视图证据显著降低候选分散性与误差,验证了该方法在精细化灾情定位中的有效性。

原文摘要 · Abstract (English)

Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.

灾情定位多模态模型地理信息应急响应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。