新基准测试发现大模型医学能力被高估,真实水平显著下降。
Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

- 构建包含1.5万+题的新波兰医学考题集,强化推理与减少猜题偏差。
- 最佳模型在新测试中表现下降28.4至31个百分点,暴露原有评估缺陷。
- 适合关注医疗AI评估可信度的研究者与开发者使用。
医学领域的大语言模型主要通过多选题问答(MCQA)评估,但该方法易受猜测策略和答案偏见影响,导致临床能力被高估。为解决这一问题,我们基于波兰医学考试构建了扩展且更具挑战性的基准,新增超过1.5万道题目、两个新领域及四项结构改进,有效降低MCQA特有误差,更真实地检验模型推理能力。我们评估了21个LLM,结果表明评估设计对性能影响显著:在更严格的设定下,最优模型(Qwen3.5-122B)在英语和波兰语考试中分别下降28.4和31个百分点。尽管缺乏数据污染证据,常规MCQA得分无法可靠反映真实医学能力。为推动后续研究,我们已公开该基准。
原文摘要 · Abstract (English)
Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategies and answer biases. To address these limitations, we introduce an expanded and more challenging benchmark based on Polish medical exams, adding over 15,000 questions, two new domains, and four structural modifications that reduce MCQA-specific artifacts and better test reasoning. We evaluate 21 LLMs and show that evaluation design strongly affects results. Under our harder setup, the best model (Qwen3.5-122B) drops by 28.4 and 31 pp on English and Polish exams, respectively. Despite low evidence of data contamination, standard MCQA scores do not reliably reflect true medical competence. To facilitate further research, we make our benchmark publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。