构建医疗LLM评估新基准,揭示安全与有效性短板
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
- 基于临床专家共识设计多维评测框架,覆盖30项关键指标
- 六款LLM平均得分57.2%,高风险场景性能下降13.3%
- 专科模型优于通用模型,尤其在安全与有效性上表现更优
大型语言模型(LLMs)在临床决策支持中潜力巨大,但其安全评估与效果验证仍面临挑战。本文构建了临床安全-有效性双轨评测基准(CSEDB),基于临床专家共识,涵盖30项关键指标,涉及危重症识别、指南依从性、用药安全等核心领域,并引入加权后果评估。32位专科医生共同开发并审核了2,069道开放式问答题,覆盖26个临床科室,模拟真实诊疗场景。对六款LLM的测试显示,整体平均得分为57.2%(安全54.7%,有效性62.3%),在高风险场景下性能显著下降13.3%(p < 0.0001)。领域专用医疗LLM相较通用模型表现更优,安全得分最高达0.912,有效性达0.861。该研究为医疗LLM临床应用提供了标准化评估工具,有助于比较分析、风险识别与优化方向制定,推动更安全有效的医疗AI部署。
原文摘要 · Abstract (English)
Large language models (LLMs) hold promise in clinical decision support but face major challenges in safety evaluation and effectiveness validation. We developed the Clinical Safety-Effectiveness Dual-Track Benchmark (CSEDB), a multidimensional framework built on clinical expert consensus, encompassing 30 criteria covering critical areas like critical illness recognition, guideline adherence, and medication safety, with weighted consequence measures. Thirty-two specialist physicians developed and reviewed 2,069 open-ended Q&A items aligned with these criteria, spanning 26 clinical departments to simulate real-world scenarios. Benchmark testing of six LLMs revealed moderate overall performance (average total score 57.2%, safety 54.7%, effectiveness 62.3%), with a significant 13.3% performance drop in high-risk scenarios (p < 0.0001). Domain-specific medical LLMs showed consistent performance advantages over general-purpose models, with relatively higher top scores in safety (0.912) and effectiveness (0.861). The findings of this study not only provide a standardized metric for evaluating the clinical application of medical LLMs, facilitating comparative analyses, risk exposure identification, and improvement directions across different scenarios, but also hold the potential to promote safer and more effective deployment of large language models in healthcare environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。