用大模型从社交媒体识别灾害影响和具体受灾地点,提升救援效率。
Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- 微调大模型识别灾害中的真实受灾地点与影响
- 在真实数据上取得0.74的受影响地点提取准确率
- 适合应急响应、灾害监测与灾后规划人员使用
大规模灾害常对人员与基础设施造成严重后果。由于大气遮挡、卫星重访周期和时间限制,依赖现场传感器、遥感图像或地理数据获取灾情信息常存在时空盲区。相比之下,社交媒体中用户发布的灾情描述可作为‘地理传感器’,提供实时位置与影响信息。但并非所有提及的地点都受灾害影响。例如,一则关于希腊马蒂火灾的推文提到‘马蒂’(受灾地)和‘希腊’‘雅典’(非受灾地)。本研究利用大语言模型(LLMs)自动识别灾害类社交媒体中提及的所有地点、灾害影响及实际受灾地点,尤其针对非正式表达、缩写和简写形式。通过微调,模型在影响提取上达到0.69的F1分数,在受影响地点提取上达0.74,显著优于预训练基线。结果表明,微调语言模型能为资源调配、态势感知和灾后恢复提供可扩展的及时决策支持。
原文摘要 · Abstract (English)
Large-scale disasters can often result in catastrophic consequences on people and infrastructure. Situation awareness about such disaster impacts generated by authoritative data from in-situ sensors, remote sensing imagery, and/or geographic data is often limited due to atmospheric opacity, satellite revisits, and time limitations. This often results in geo-temporal information gaps. In contrast, impact-related social media posts can act as "geo-sensors" during a disaster, where people describe specific impacts and locations. However, not all locations mentioned in disaster-related social media posts relate to an impact. Only the impacted locations are critical for directing resources effectively. e.g., "The death toll from a fire which ripped through the Greek coastal town of #Mati stood at 80, with dozens of people unaccounted for as forensic experts tried to identify victims who were burned alive #Greecefires #AthensFires #Athens #Greece." contains impacted location "Mati" and non-impacted locations "Greece" and "Athens". This research uses Large Language Models (LLMs) to identify all locations, impacts and impacted locations mentioned in disaster-related social media posts. In the process, LLMs are fine-tuned to identify only impacts and impacted locations (as distinct from other, non-impacted locations), including locations mentioned in informal expressions, abbreviations, and short forms. Our fine-tuned model demonstrates efficacy, achieving an F1-score of 0.69 for impact and 0.74 for impacted location extraction, substantially outperforming the pre-trained baseline. These robust results confirm the potential of fine-tuned language models to offer a scalable solution for timely decision-making in resource allocation, situational awareness, and post-disaster recovery planning for responders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。