用大模型自动解析灾害文本位置,精准定位到县市级。
Subnational Geocoding of Global Disasters Using Large Language Models
- 用GPT-4o处理文本位置,交叉核对三大地理数据库
- 为14,215个灾害事件生成县市级地理坐标,覆盖17,948个地点
- 无需人工干预,支持多源验证,适合灾害分析与空间研究
灾害事件的次国家级地理位置数据对风险评估和减灾至关重要。现有灾害数据库如EM-DAT常以非结构化文本形式报告位置,存在粒度不一、拼写差异等问题,难以与空间数据集整合。本文提出一种全自动的LLM辅助工作流,利用GPT-4o处理并清洗文本位置信息,并通过交叉核对GADM、OpenStreetMap和Wikidata三个独立地理信息库来分配几何坐标。根据各源的一致性与可用性,为每个位置赋予可靠性评分。该方法应用于2000至2024年的EM-DAT数据集,成功为14,215个事件、17,948个唯一地点生成次国家级几何信息。相比以往方法,本方案无需人工干预,覆盖所有灾害类型,支持多源交叉验证,并可灵活映射至不同地理框架。此外,研究展示了大模型从非结构化文本中提取与结构化地理信息的潜力,为相关分析提供可扩展、可靠的解决方案。
原文摘要 · Abstract (English)
Subnational location data of disaster events are critical for risk assessment and disaster risk reduction. Disaster databases such as EM-DAT often report locations in unstructured textual form, with inconsistent granularity or spelling, that make it difficult to integrate with spatial datasets. We present a fully automated LLM-assisted workflow that processes and cleans textual location information using GPT-4o, and assigns geometries by cross-checking three independent geoinformation repositories: GADM, OpenStreetMap and Wikidata. Based on the agreement and availability of these sources, we assign a reliability score to each location while generating subnational geometries. Applied to the EM-DAT dataset from 2000 to 2024, the workflow geocodes 14,215 events across 17,948 unique locations. Unlike previous methods, our approach requires no manual intervention, covers all disaster types, enables cross-verification across multiple sources, and allows flexible remapping to preferred frameworks. Beyond the dataset, we demonstrate the potential of LLMs to extract and structure geographic information from unstructured text, offering a scalable and reliable method for related analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。