arXiv:2503.10694cs.CL2025-03ICML被引 45

医学大模型评测需验证真实能力,而非仅刷榜单。

Medical Large Language Model Benchmarks Should Prioritize Construct Validity

  • 用心理测量学的建构效度评估医疗大模型基准测试
  • 实证发现主流评测存在显著效度缺口
  • 适合关注模型真实临床能力的研究者

医学大语言模型研究常宣称具备临床知识编码与医生推理能力,这些主张多依赖于竞赛性质的基准测试,其构建方式沿袭自主流机器学习领域,通常采用医学执照考试题目。然而,若要真实衡量进展,这些基准必须准确反映其所代表的实际临床任务。本文提出,医学大模型基准测试应具备可验证的建构效度——即能否有效测量其理论目标。借鉴心理学测试中的建构效度概念,我们论证了该框架在模型评估中的适用性,并通过真实临床数据开展概念验证实验,揭示了现有主流基准在建构效度上的显著不足。最后,本文展望未来应建立以有效基准为核心的新型医学大模型评估生态。

原文摘要 · Abstract (English)

Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks; a tradition inherited from mainstream machine learning. But how do we separate real progress from a leaderboard flex? Medical LLM benchmarks, much like those in other fields, are arbitrarily constructed using medical licensing exam questions. For these benchmarks to truly measure progress, they must accurately capture the real-world tasks they aim to represent. In this position paper, we argue that medical LLM benchmarks should (and indeed can) be empirically evaluated for their construct validity. In the psychological testing literature, "construct validity" refers to the ability of a test to measure an underlying "construct", that is the actual conceptual target of evaluation. By drawing an analogy between LLM benchmarks and psychological tests, we explain how frameworks from this field can provide empirical foundations for validating benchmarks. To put these ideas into practice, we use real-world clinical data in proof-of-concept experiments to evaluate popular medical LLM benchmarks and report significant gaps in their construct validity. Finally, we outline a vision for a new ecosystem of medical LLM evaluation centered around the creation of valid benchmarks.

医学AI模型评估建构效度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。