arXiv:2505.14107cs.CLcs.AI2025-05ACL被引 17

新基准DiagnosisArena测试大模型临床诊断能力,发现顶尖模型准确率不足五成。

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

  • 构建跨28个专科的1113对病案-诊断数据集,严格防泄露。
  • 最强模型o3仅达51.12%准确率,暴露诊断推理瓶颈。
  • 适合医疗AI研究者评估模型真实临床推理水平。

具备复杂推理能力的大语言模型在应对科学挑战(如复杂临床场景)方面前景广阔。为确保其在真实医疗环境中的安全有效应用,亟需系统性地评估当前模型的诊断能力。鉴于现有医学基准在评估高级诊断推理方面的局限性,我们提出DiagnosisArena——一个全面且具有挑战性的基准,用于严格评估专业级诊断能力。DiagnosisArena包含1,113对分段患者病例与对应诊断,覆盖28个医学专科,数据源自10本顶级医学期刊发表的临床病例报告。该基准通过多轮筛选与评审流程构建,结合人工智能系统与人类专家,进行彻底检查以防止数据泄露。研究发现,即使是最先进的推理模型o3、o1和DeepSeek-R1,准确率也仅为51.12%、31.09%和17.79%。这一结果凸显了当前大语言模型在面对临床诊断推理挑战时存在显著泛化瓶颈。DiagnosisArena旨在推动AI诊断推理能力的进一步发展,为解决真实世界临床诊断难题提供更有效的方案。相关基准与评估工具已开源:https://github.com/SPIRAL-MED/DiagnosisArena。

原文摘要 · Abstract (English)

The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable their safe and effective deployment in real-world healthcare settings, it is urgently necessary to benchmark the diagnostic capabilities of current models systematically. Given the limitations of existing medical benchmarks in evaluating advanced diagnostic reasoning, we present DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess professional-level diagnostic competence. DiagnosisArena consists of 1,113 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, deriving from clinical case reports published in 10 top-tier medical journals. The benchmark is developed through a meticulous construction pipeline, involving multiple rounds of screening and review by both AI systems and human experts, with thorough checks conducted to prevent data leakage. Our study reveals that even the most advanced reasoning models, o3, o1, and DeepSeek-R1, achieve only 51.12%, 31.09%, and 17.79% accuracy, respectively. This finding highlights a significant generalization bottleneck in current large language models when faced with clinical diagnostic reasoning challenges. Through DiagnosisArena, we aim to drive further advancements in AI's diagnostic reasoning capabilities, enabling more effective solutions for real-world clinical diagnostic challenges. We provide the benchmark and evaluation tools for further research and development https://github.com/SPIRAL-MED/DiagnosisArena.

临床诊断大模型评测医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。