研究发现大模型在选择‘以上都不对’时表现大幅下降,暴露其缺乏判断不确定性的能力。
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
- 通过28个大模型在MMLU上的实验,分析'以上都不对'选项对性能的影响。
- 当正确答案是'以上都不对'时,模型准确率普遍下降30%-50%,尤其在伦理类任务中降幅达48.1%。
- 揭示大模型在不确定性判断上的缺陷,对评测基准设计和真实应用有重要启示。
多项选择题中包含'以上都不对'(NA)选项被广泛用于教育评估,被认为能更有效检验真实知识。然而,其对大语言模型(LLMs)评估的影响仍不明确。我们通过对28个LLMs在MMLU基准上的系统性实验,研究了NA选项对模型性能与置信度校准的影响。分析显示,当正确答案为NA时,所有模型性能均出现30%-50%的持续下降,无论模型规模如何——表明大模型缺乏系统评估并排除所有选项的能力。该退化现象具有显著领域依赖性:数学推理任务仅下降14.6%,但在需处理不确定性的商业伦理任务中下降高达48.1%。结果对评测基准设计提出重要启示,并引发对大模型在现实应用中应对不确定性的质疑。
原文摘要 · Abstract (English)
Multiple-choice exam questions with "None of the above" (NA) options have been extensively studied in educational testing, in which existing research suggests that they better assess true knowledge. However, their impact on Large Language Models (LLMs) evaluation remains underexplored. Through systematic experiments with 28 LLMs on the MMLU benchmark, we examine how NA options affect model performance and confidence calibration. Our analysis reveals that NA options, when used as the correct answer, lead to a consistent 30-50\% performance drop across models regardless of scale--suggesting that LLMs lack the meta-cognitive ability to systematically evaluate and reject all given options when none are correct. This degradation shows strong domain dependence, with minimal impact on mathematical reasoning (14.6\% drop) but severe effects on tasks requiring uncertainty handling like business ethics (48.1\% drop). Our results highlight important implications for benchmark design and raise questions about LLMs' ability to handle uncertainty in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。