arXiv:2503.11348cs.CL2025-03被引 3

评测大模型在灾难场景下的常识推理能力,发现顶尖模型准确率仅37%。

RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation

  • 构建包含1789个实例的灾难常识推理数据集,覆盖多时间阶段响应需求。
  • 顶尖模型如GPT-4在即时应对建议上人类评估准确率仅37%。
  • 适合关注大模型应急管理能力与可信推理的研究者使用。

当人们面临自然灾害时,会产生一类特殊的常识推理问题。为研究该主题,我们提出 extsf{RESPONSE},一个由人工标注的数据集,包含1789个标注实例和6037组问题,用于评估大语言模型(LLM)在不同时间阶段的灾难情境下进行常识推理的能力。数据集涵盖事件描述、缺失资源、时间敏感性解决方案及其理由,并有环境工程师验证子集。通过自动指标与人工评估,我们对比了模型生成建议与人类回答的表现。结果表明,即使最先进的模型如GPT-4,在即时响应行动的人类评估中也仅达到37%的正确率,凸显大模型在危机情境下常识推理能力仍有巨大提升空间。

原文摘要 · Abstract (English)

An interesting class of commonsense reasoning problems arises when people are faced with natural disasters. To investigate this topic, we present \textsf{RESPONSE}, a human-curated dataset containing 1789 annotated instances featuring 6037 sets of questions designed to assess LLMs' commonsense reasoning in disaster situations across different time frames. The dataset includes problem descriptions, missing resources, time-sensitive solutions, and their justifications, with a subset validated by environmental engineers. Through both automatic metrics and human evaluation, we compare LLM-generated recommendations against human responses. Our findings show that even state-of-the-art models like GPT-4 achieve only 37\% human-evaluated correctness for immediate response actions, highlighting significant room for improvement in LLMs' ability for commonsense reasoning in crises.

常识推理灾难应对大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。