arXiv:2505.23802cs.CLcs.AI2025-05被引 75

医学大模型需真实临床任务评估,此研究构建了全面测评框架。

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

  • 基于29位医生制定5类22子类121项任务的医学评估分类体系
  • 覆盖35个基准测试,发现顶尖模型在病历生成上准确率达0.85
  • 提出LLM评审法,成本低且与医生评分高度一致

尽管大语言模型在医学执照考试中接近满分,但此类评估无法反映真实临床实践的复杂性。本文提出MedHELM,一个可扩展的医学任务评估框架,包含三大贡献:第一,由29名临床医生验证的分类体系,涵盖5个类别、22个子类别和121项任务;第二,包含35个基准测试(17个现有,18个新设计)的综合评测套件,覆盖全部分类;第三,采用改进的评估方法(使用LLM-jury)并进行成本-性能分析。对9个前沿模型的评估显示显著性能差异:高级推理模型(DeepSeek R1:66%胜率;o3-mini:64%胜率)表现优异,而Claude 3.5 Sonnet以40%更低计算成本达到相近效果。在标准化准确率(0-1)下,多数模型在病历生成(0.73-0.85)和患者沟通教育(0.78-0.83)表现良好,医学研究辅助(0.65-0.75)中等,临床决策支持(0.56-0.72)和行政流程(0.53-0.63)普遍偏低。所提LLM-jury方法与医生评分一致性(ICC=0.47)优于医生间一致性(ICC=0.43)及自动基线(ROUGE-L: 0.36, BERTScore-F1: 0.44)。结果强调了真实任务评估的重要性,并开源了该框架。

原文摘要 · Abstract (English)

While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice. We introduce MedHELM, an extensible evaluation framework for assessing LLM performance for medical tasks with three key contributions. First, a clinician-validated taxonomy spanning 5 categories, 22 subcategories, and 121 tasks developed with 29 clinicians. Second, a comprehensive benchmark suite comprising 35 benchmarks (17 existing, 18 newly formulated) providing complete coverage of all categories and subcategories in the taxonomy. Third, a systematic comparison of LLMs with improved evaluation methods (using an LLM-jury) and a cost-performance analysis. Evaluation of 9 frontier LLMs, using the 35 benchmarks, revealed significant performance variation. Advanced reasoning models (DeepSeek R1: 66% win-rate; o3-mini: 64% win-rate) demonstrated superior performance, though Claude 3.5 Sonnet achieved comparable results at 40% lower estimated computational cost. On a normalized accuracy scale (0-1), most models performed strongly in Clinical Note Generation (0.73-0.85) and Patient Communication & Education (0.78-0.83), moderately in Medical Research Assistance (0.65-0.75), and generally lower in Clinical Decision Support (0.56-0.72) and Administration & Workflow (0.53-0.63). Our LLM-jury evaluation method achieved good agreement with clinician ratings (ICC = 0.47), surpassing both average clinician-clinician agreement (ICC = 0.43) and automated baselines including ROUGE-L (0.36) and BERTScore-F1 (0.44). Claude 3.5 Sonnet achieved comparable performance to top models at lower estimated cost. These findings highlight the importance of real-world, task-specific evaluation for medical use of LLMs and provides an open source framework to enable this.

医学AI大模型评估临床应用基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。