arXiv:2602.08872cs.CLcs.IR2026-02

用大模型提升人道危机文本的地理信息提取公平性

Large Language Models for Geolocation Extraction in Humanitarian Crisis Response

  • 两步法:少样本大模型识别地名,再用智能体结合上下文消歧
  • 在低曝光地区地名提取准确率显著提升,公平性指标改善明显
  • 适合关注人道主义数据公平性的研究者与应急响应系统开发者

人道危机需要及时准确的地理信息以支持有效响应。然而,现有自动文本地理信息提取系统常延续既有地理与社会经济偏见,导致受影响区域可见性不均。本文探讨大语言模型(LLMs)是否能缓解此类地理差异。提出一种两步框架:先用少量样本的LLM进行命名实体识别,再通过基于代理的地理编码模块利用上下文解决地名歧义。在扩展版HumSet数据集上,对齐预训练模型和规则系统进行了准确性与公平性评估。结果表明,基于LLM的方法显著提升了人道文本中地理信息提取的精度与公平性,尤其在代表性不足地区表现更优。该工作融合大模型推理能力与负责任、包容性AI原则,推动更具公平性的危机响应地理空间数据体系,助力实现‘不遗漏任何地区’的危机分析目标。

原文摘要 · Abstract (English)

Humanitarian crises demand timely and accurate geographic information to inform effective response efforts. Yet, automated systems that extract locations from text often reproduce existing geographic and socioeconomic biases, leading to uneven visibility of crisis-affected regions. This paper investigates whether Large Language Models (LLMs) can address these geographic disparities in extracting location information from humanitarian documents. We introduce a two-step framework that combines few-shot LLM-based named entity recognition with an agent-based geocoding module that leverages context to resolve ambiguous toponyms. We benchmark our approach against state-of-the-art pretrained and rule-based systems using both accuracy and fairness metrics across geographic and socioeconomic dimensions. Our evaluation uses an extended version of the HumSet dataset with refined literal toponym annotations. Results show that LLM-based methods substantially improve both the precision and fairness of geolocation extraction from humanitarian texts, particularly for underrepresented regions. By bridging advances in LLM reasoning with principles of responsible and inclusive AI, this work contributes to more equitable geospatial data systems for humanitarian response, advancing the goal of leaving no place behind in crisis analytics.

地理信息大模型人道救援公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。