arXiv:2602.16747cs.LGcs.AI2026-02被引 3

LiveClin构建动态临床评估基准,避免数据污染和过时问题。

LiveClin: A Live Clinical Benchmark without Leakage

  • 基于真实病例与医生协作,生成多模态临床评估场景
  • 26个模型最高准确率仅35.7%,远低于主治医师水平
  • 适合关注医疗大模型真实能力评估的研究者使用

医学大模型评估的可靠性因数据污染和知识过时而严重受损,导致静态基准分数虚高。为解决这一问题,我们提出LiveClin——一个模拟真实临床实践的动态评估基准。其数据源自最新同行评审病例报告,每两年更新一次,确保临床时效性并抵御数据污染。通过239名医师验证的AI-人工协作流程,我们将真实患者案例转化为涵盖完整诊疗路径的复杂多模态评估任务。当前基准包含1,407篇病例报告和6,605个问题。对26个模型的评估显示,这些真实场景极具挑战性,表现最佳模型的病例准确率仅为35.7%。与人类专家对比,主任医师准确率最高,住院医师紧随其后,两者均显著超越多数模型。LiveClin因此提供了一个持续演进、临床扎根的框架,助力医疗大模型缩小差距,提升真实可用性。数据与代码已公开于https://github.com/AQ-MedAI/LiveClin。

原文摘要 · Abstract (English)

The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address these challenges, we introduce LiveClin, a live benchmark designed for approximating real-world clinical practice. Built from contemporary, peer-reviewed case reports and updated biannually, LiveClin ensures clinical currency and resists data contamination. Using a verified AI-human workflow involving 239 physicians, we transform authentic patient cases into complex, multimodal evaluation scenarios that span the entire clinical pathway. The benchmark currently comprises 1,407 case reports and 6,605 questions. Our evaluation of 26 models on LiveClin reveals the profound difficulty of these real-world scenarios, with the top-performing model achieving a Case Accuracy of just 35.7%. In benchmarking against human experts, Chief Physicians achieved the highest accuracy, followed closely by Attending Physicians, with both surpassing most models. LiveClin thus provides a continuously evolving, clinically grounded framework to guide the development of medical LLMs towards closing this gap and achieving greater reliability and real-world utility. Our data and code are publicly available at https://github.com/AQ-MedAI/LiveClin.

医疗大模型评估基准临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。