构建实时更新的医疗大模型评测基准,避免数据污染并精准评估临床推理能力。
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation
- 每周从真实医患社区抓取病例,确保训练与测试数据时间隔离。
- 用16,702条细粒度评分标准评估模型,准确率仅39.2%且84%模型在新病例上退化。
- 发现模型失败主因是无法结合患者具体情境应用知识,非事实记忆错误。
大语言模型(LLMs)在高风险临床场景中的部署需要严格可靠的评估。然而,现有医学评测基准仍为静态,存在两大关键缺陷:(1) 数据污染——测试集意外泄露至训练语料中,导致性能虚高;(2) 时间错配——无法反映医学知识的快速演进。此外,当前开放问答型临床推理评估多依赖浅层词法重合(如ROUGE)或主观的LLM作为裁判评分,均难以验证临床正确性。为此,我们提出LiveMedBench,一个持续更新、无数据污染、基于评分标准的评测基准,每周从在线医学社区采集真实临床案例,确保与模型训练数据严格时序分离。我们设计了多智能体临床数据净化框架,过滤原始噪声并依据循证医学原则验证临床完整性。评估方面,开发自动化评分框架,将医生回答拆解为细粒度、案例特异的评判维度,其与专家医师的一致性显著优于LLM-as-a-Judge。截至目前,LiveMedBench包含2,756个真实病例,覆盖38个医学专科及多种语言,配有16,702个独特评价标准。对38个LLMs的全面评估显示,最优模型准确率仅为39.2%,84%模型在截止日期后的病例上表现下降,证实普遍存在的数据污染风险。错误分析进一步指出,主要瓶颈在于上下文应用而非事实知识,35%-48%的失败源于无法将医学知识适配至患者特定约束。
原文摘要 · Abstract (English)
The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffering from two critical limitations: (1) data contamination, where test sets inadvertently leak into training corpora, leading to inflated performance estimates; and (2) temporal misalignment, failing to capture the rapid evolution of medical knowledge. Furthermore, current evaluation metrics for open-ended clinical reasoning often rely on either shallow lexical overlap (e.g., ROUGE) or subjective LLM-as-a-Judge scoring, both inadequate for verifying clinical correctness. To bridge these gaps, we introduce LiveMedBench, a continuously updated, contamination-free, and rubric-based benchmark that weekly harvests real-world clinical cases from online medical communities, ensuring strict temporal separation from model training data. We propose a Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles. For evaluation, we develop an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge. To date, LiveMedBench comprises 2,756 real-world cases spanning 38 medical specialties and multiple languages, paired with 16,702 unique evaluation criteria. Extensive evaluation of 38 LLMs reveals that even the best-performing model achieves only 39.2%, and 84% of models exhibit performance degradation on post-cutoff cases, confirming pervasive data contamination risks. Error analysis further identifies contextual application-not factual knowledge-as the dominant bottleneck, with 35-48% of failures stemming from the inability to tailor medical knowledge to patient-specific constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。