arXiv:2410.20707cs.CL2024-10被引 8

构建灾难响应评测基准,检验大模型在灾情决策中的表现

DisasterQA: A Benchmark for Assessing the performance of LLMs in Disaster Response

  • 从六个在线源构建灾难响应问答数据集
  • 五款大模型平均准确率不足60%,需提升知识能力
  • 适合研究灾难智能系统与应急决策的开发者

灾难可能导致大量人员伤亡,快速响应至关重要。大型语言模型(LLMs)在处理海量文本信息方面表现出色,可为灾情提供情境理解。然而,它们是否适合用于灾情建议与决策仍存疑问。为此,我们基于六个在线来源构建了名为DisasterQA的评测基准,涵盖广泛灾情应对主题。评估了五款主流大模型,每款采用四种提示策略,在该基准上测量准确率与置信度(通过Logprobs)。结果表明,当前大模型在灾情知识方面仍需改进。本研究旨在推动大模型在灾情响应领域的进一步发展,使其未来能与应急管理人员协同工作。

原文摘要 · Abstract (English)

Disasters can result in the deaths of many, making quick response times vital. Large Language Models (LLMs) have emerged as valuable in the field. LLMs can be used to process vast amounts of textual information quickly providing situational context during a disaster. However, the question remains whether LLMs should be used for advice and decision making in a disaster. To evaluate the capabilities of LLMs in disaster response knowledge, we introduce a benchmark: DisasterQA created from six online sources. The benchmark covers a wide range of disaster response topics. We evaluated five LLMs each with four different prompting methods on our benchmark, measuring both accuracy and confidence levels through Logprobs. The results indicate that LLMs require improvement on disaster response knowledge. We hope that this benchmark pushes forth further development of LLMs in disaster response, ultimately enabling these models to work alongside. emergency managers in disasters.

灾难响应大模型评测LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。