首个面向临床场景的LLM评估基准,专为真实医疗应用设计。
MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
- 构建涵盖64个专科的22个中文临床数据集,覆盖五维临床能力。
- 顶尖模型在结构化任务准确率达87.8%,但患者交互场景仅61.3%。
- 采用医生标注的专用评判模型,确保评估结果可复现且贴近临床需求。
大型语言模型(LLMs)在医疗领域具革命性潜力,但实际部署受限于缺乏评估真实临床价值的框架。现有基准多测试静态知识,无法反映临床实践所需的动态、应用导向能力。为此,我们提出医学LLM基准 MLB,全面评估模型在基础医学知识与情景推理两方面的表现。MLB围绕五大维度展开:医学知识(MedKQA)、安全与伦理(MedSE)、病历理解(MedRU)、智能服务(SmartServ)和智慧医疗(SmartCare)。该基准整合了22个数据集(其中17个为新构建),源自多样化的中国临床来源,覆盖64个临床专科。其设计包含由300名持证医师参与的严格数据筛选流程。此外,我们提出一种可扩展的评估方法,基于19,000条专家标注数据进行监督微调(SFT)训练的专用评判模型。对10个主流模型的综合评估显示显著转化鸿沟:排名第一的Kimi-K2-Instruct(总体准确率77.3%)在信息抽取等结构化任务中达87.8%准确率,但在患者交互场景中下降至61.3%。而小型模型Baichuan-M2-32B在安全评估中表现优异(90.6%),表明针对性训练同样关键。所提出的评判模型在人类-人工智能一致性上达到92.1%准确率、94.37% F1值与81.3% Cohen's Kappa,验证了可重复且专家对齐的评估协议。MLB为开发具备临床实用性的LLMs提供了严谨框架。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) presents transformative potential for healthcare, yet practical deployment is hindered by the absence of frameworks that assess real-world clinical utility. Existing benchmarks test static knowledge, failing to capture the dynamic, application-oriented capabilities required in clinical practice. To bridge this gap, we introduce a Medical LLM Benchmark MLB, a comprehensive benchmark evaluating LLMs on both foundational knowledge and scenario-based reasoning. MLB is structured around five core dimensions: Medical Knowledge (MedKQA), Safety and Ethics (MedSE), Medical Record Understanding (MedRU), Smart Services (SmartServ), and Smart Healthcare (SmartCare). The benchmark integrates 22 datasets (17 newly curated) from diverse Chinese clinical sources, covering 64 clinical specialties. Its design features a rigorous curation pipeline involving 300 licensed physicians. Besides, we provide a scalable evaluation methodology, centered on a specialized judge model trained via Supervised Fine-Tuning (SFT) on expert annotations. Our comprehensive evaluation of 10 leading models reveals a critical translational gap: while the top-ranked model, Kimi-K2-Instruct (77.3% accuracy overall), excels in structured tasks like information extraction (87.8% accuracy in MedRU), performance plummets in patient-facing scenarios (61.3% in SmartServ). Moreover, the exceptional safety score (90.6% in MedSE) of the much smaller Baichuan-M2-32B highlights that targeted training is equally critical. Our specialized judge model, trained via SFT on a 19k expert-annotated medical dataset, achieves 92.1% accuracy, an F1-score of 94.37%, and a Cohen's Kappa of 81.3% for human-AI consistency, validating a reproducible and expert-aligned evaluation protocol. MLB thus provides a rigorous framework to guide the development of clinically viable LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。