构建灾难管理问答新基准,测试模型在混乱信息下的真实推理能力
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management
- 通过人机协作构建3000个跨8类灾难的问答数据集
- 开放题采用关键点验证法,强调事实完整而非冗长回答
- 发现多数模型在真实噪声环境下性能急剧下降,暴露可靠性短板
灾难管理中的精准问答需处理不确定和冲突的信息,而现有基准多基于清晰证据,难以反映真实场景。我们提出DisastQA,一个包含3000个经严格验证的问题(2000个选择题,1000个开放题)的大规模基准,覆盖八类灾难。数据通过人机协作与分层采样构建,确保覆盖均衡。模型在从闭卷到噪声证据融合的不同证据条件下评估,可分离内部知识与在不完善信息下的推理能力。针对开放题,我们提出基于人工验证的关键点评估协议,强调事实完整性而非篇幅。20个模型的实验显示,其表现与通用排行榜(如MMLU-Pro)存在显著差异:尽管近期开源模型在干净环境下接近商用系统,但在现实噪声下性能骤降,暴露出灾难响应中的关键可靠性缺陷。所有代码、数据与评估资源已公开于https://github.com/TamuChen18/DisastQA_open。
原文摘要 · Abstract (English)
Accurate question answering (QA) in disaster management requires reasoning over uncertain and conflicting information, a setting poorly captured by existing benchmarks built on clean evidence. We introduce DisastQA, a large-scale benchmark of 3,000 rigorously verified questions (2,000 multiple-choice and 1,000 open-ended) spanning eight disaster types. The benchmark is constructed via a human-LLM collaboration pipeline with stratified sampling to ensure balanced coverage. Models are evaluated under varying evidence conditions, from closed-book to noisy evidence integration, enabling separation of internal knowledge from reasoning under imperfect information. For open-ended QA, we propose a human-verified keypoint-based evaluation protocol emphasizing factual completeness over verbosity. Experiments with 20 models reveal substantial divergences from general-purpose leaderboards such as MMLU-Pro. While recent open-weight models approach proprietary systems in clean settings, performance degrades sharply under realistic noise, exposing critical reliability gaps for disaster response. All code, data, and evaluation resources are available at https://github.com/TamuChen18/DisastQA_open.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。