arXiv:2502.14127cs.CL2025-02ACL被引 63

指出大模型多选题评测的四大缺陷并提出改进方案

Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above

  • 用生成式答题模拟人类测试,更贴近真实使用场景
  • 多选题数据集存在泄露、不可答、捷径等严重问题
  • 借鉴教育测评方法优化题目设计与评分机制

多选题问答(MCQA)因形式简单且类人测试而广泛用于大模型评估,但本文指出其存在根本性缺陷:难以检验生成能力与主观判断;与实际应用需求不匹配;无法全面考察知识掌握。为此,建议采用基于人类测试的生成式评测范式,让大模型自主构建并解释答案,更准确反映用户需求与知识水平,同时保持可评分性。进一步发现,即便在适用场景下,现有MCQA数据集仍存在答案泄露、不可回答、答题捷径和题目饱和等问题。针对每类问题,本文提出教育学中的解决方案,如使用评分标准指导题目设计、引入打分机制抑制猜测行为、利用项目反应理论构造更难题目。最后讨论了大模型在多选题中的错误、鲁棒性、偏见及不忠实解释,并证明前述优化措施能更好衡量或缓解这些问题。本文主张不应放弃多选题,而是应基于教育测评原理进行系统性改进。

原文摘要 · Abstract (English)

Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform. We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity; 2) match LLM use cases; and 3) fully test knowledge. We instead advocate for generative formats based on human testing, where LLMs construct and explain answers, better capturing user needs and knowledge while remaining easy to score. We then show even when MCQA is a useful format, its datasets suffer from: leakage; unanswerability; shortcuts; and saturation. In each issue, we give fixes from education, like rubrics to guide MCQ writing; scoring methods to bridle guessing; and Item Response Theory to build harder MCQs. Lastly, we discuss LLM errors in MCQA, robustness, biases, and unfaithful explanations, showing how our prior solutions better measure or address these issues. While we do not need to desert MCQA, we encourage more efforts in refining the task based on educational testing, advancing evaluations.

评测方法大模型评估多选题教育测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。