LLM在医学多选题中表现好,但换成自由作答后成绩暴跌,暴露了测试题设计的漏洞。
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
- 用配对的自由作答题与多选题对比,检验LLM真实医学能力
- LLM自由作答准确率比多选题平均下降39.43%,超人类22.29%的降幅
- 多选题格式本身让模型靠猜得分,适合用于评估人机问答的改进方法
大型语言模型(LLMs)在多选题(MCQ)基准上的表现常被视作其医学能力的证明。我们假设,这种表现可能部分是虚假的,受非医学知识与推理能力因素影响。为此,我们构建了一个包含自由作答题与配对多选题的新基准(FreeMedQA),评估了三种先进模型(GPT-4o、GPT-3.5、LLama-3-70B-instruct)。结果显示,模型在自由作答题上平均准确率比多选题下降39.43%(p = 1.3 × 10⁻⁵),高于人类下降的22.29%。通过逐步遮蔽题干的掩码实验发现,当题干完全遮蔽时,多选题平均正确率仅比随机猜测高6.70%(p = 0.002),其中GPT-4o最高达37.34%;而所有模型在自由作答中表现接近零。结果表明,现有医学多选题基准容易夸大模型能力,且提示可借助模型评估自由作答以改进人机评估体系。
原文摘要 · Abstract (English)
The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10-5) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。