arXiv:2505.20249cs.CLcs.AI2025-05ACL被引 1

首个评估大模型理解极端天气影响能力的基准,助力气候适应研究。

WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models

  • 构建四阶段流程的数据集,整合地方报纸中的灾后应对记录。
  • 设计多标签分类与排序问答两项任务,验证模型在灾害影响理解上的表现。
  • 为气候适应系统开发提供可复用数据与评估框架,适合气候智能研究者使用。

气候变化适应需要理解极端天气对社会的影响,而大语言模型(LLMs)在此领域应用潜力尚未充分探索,主要受限于高质量语料库收集困难及缺乏可用基准。区域报纸中保存的气候事件记录了社区在灾后的适应与恢复过程,但原始语料处理复杂。本研究首先通过四阶段精心设计的构建流程,开发了一个极端天气影响数据集;随后提出 WXImpactBench,这是首个用于评估大模型在极端天气影响理解方面能力的基准。该基准包含两项评估任务:多标签分类与基于排序的问答。对一系列大模型的广泛实验提供了关于发展极端天气影响理解与气候适应系统所面临挑战的一手分析。所构建的数据集及评估框架代码已公开,旨在帮助社会增强对灾害脆弱性的防护能力。

原文摘要 · Abstract (English)

Climate change adaptation requires the understanding of disruptive weather impacts on society, where large language models (LLMs) might be applicable. However, their effectiveness is under-explored due to the difficulty of high-quality corpus collection and the lack of available benchmarks. The climate-related events stored in regional newspapers record how communities adapted and recovered from disasters. However, the processing of the original corpus is non-trivial. In this study, we first develop a disruptive weather impact dataset with a four-stage well-crafted construction pipeline. Then, we propose WXImpactBench, the first benchmark for evaluating the capacity of LLMs on disruptive weather impacts. The benchmark involves two evaluation tasks, multi-label classification and ranking-based question answering. Extensive experiments on evaluating a set of LLMs provide first-hand analysis of the challenges in developing disruptive weather impact understanding and climate change adaptation systems. The constructed dataset and the code for the evaluation framework are available to help society protect against vulnerabilities from disasters.

大模型评测气候适应极端天气社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。