首个医疗大模型评估框架,揭示现有基准的临床脱节与数据隐患
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
- 构建生命周期评估框架,分五个阶段系统审查医疗基准设计
- 53个基准中普遍存临床偏离、数据污染与安全评估缺失问题
- 适合医疗AI研究者、审评机构及政策制定者参考使用
大语言模型在医疗领域潜力巨大,催生了众多评估基准。但这些基准普遍存在临床真实性不足、数据管理不健全、安全评估指标缺失等问题。为此,我们提出首个面向医疗基准的生命周期评估框架MedCheck,将基准开发拆解为设计、数据、评估、部署、治理五个连续阶段,并提供46项医学定制化检查标准。基于MedCheck对53个医疗大模型基准进行实证分析,发现普遍存在与临床实践脱节、数据污染风险未管控、模型鲁棒性与不确定性感知等安全维度被忽视等系统性缺陷。MedCheck既可诊断现有基准问题,也可作为推动医疗AI评估标准化、可靠化与透明化的行动指南。
原文摘要 · Abstract (English)
Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However, concerns persist regarding the reliability of these benchmarks, which often lack clinical fidelity, robust data management, and safety-oriented evaluation metrics. To address these shortcomings, we introduce MedCheck, the first lifecycle-oriented assessment framework specifically designed for medical benchmarks. Our framework deconstructs a benchmark's development into five continuous stages, from design to governance, and provides a comprehensive checklist of 46 medically-tailored criteria. Using MedCheck, we conducted an in-depth empirical evaluation of 53 medical LLM benchmarks. Our analysis uncovers widespread, systemic issues, including a profound disconnect from clinical practice, a crisis of data integrity due to unmitigated contamination risks, and a systematic neglect of safety-critical evaluation dimensions like model robustness and uncertainty awareness. Based on these findings, MedCheck serves as both a diagnostic tool for existing benchmarks and an actionable guideline to foster a more standardized, reliable, and transparent approach to evaluating AI in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。