arXiv:2503.14996cs.CL2025-03ACL被引 38

揭示大模型多选题评估中的评分不一致问题,指出当前方法可能低估模型真实能力。

Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

  • 分析不同答案提取方法与人类判断的一致性
  • 发现传统评估常低估模型表现,基于LLM的提取器有系统性错误
  • 提示需统一评估标准,适合关注评测可靠性的研究者

大语言模型(LLMs)最常用的评估任务之一是多选题问答(MCQA)。尽管开放题评估困难,但理论上MCQA因答案可直接比对预设选项而更易评估。然而近期研究质疑其可靠性,指出多种因素会显著影响模型性能报告,尤其当模型先生成自由文本再选择答案时。本文系统分析现有答案提取方法是否符合人类判断,并考察提示中答案约束在不同领域的影响。实验表明,传统评估策略常低估LLM能力,而基于LLM的答案提取器存在系统性偏差。此外,我们揭示了在提示中加入格式约束以简化提取与允许自由生成以提升推理之间的根本权衡。研究呼吁建立标准化评估方法,强调需更可靠、一致的MCQA评估实践。

原文摘要 · Abstract (English)

One of the most widely used tasks for evaluating Large Language Models (LLMs) is Multiple-Choice Question Answering (MCQA). While open-ended question answering tasks are more challenging to evaluate, MCQA tasks are, in principle, easier to assess, as the model's answer is thought to be simple to extract and is compared directly to a set of predefined choices. However, recent studies have started to question the reliability of MCQA evaluation, showing that multiple factors can significantly impact the reported performance of LLMs, especially when the model generates free-form text before selecting one of the answer choices. In this work, we shed light on the inconsistencies of MCQA evaluation strategies, which can lead to inaccurate and misleading model comparisons. We systematically analyze whether existing answer extraction methods are aligned with human judgment, and how they are influenced by answer constraints in the prompt across different domains. Our experiments demonstrate that traditional evaluation strategies often underestimate LLM capabilities, while LLM-based answer extractors are prone to systematic errors. Moreover, we reveal a fundamental trade-off between including format constraints in the prompt to simplify answer extraction and allowing models to generate free-form text to improve reasoning. Our findings call for standardized evaluation methodologies and highlight the need for more reliable and consistent MCQA evaluation practices.

大模型评估多选题评测一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。