arXiv:2412.10849cs.AIcs.CL2024-12被引 82

大语言模型在临床推理任务中表现超越医生,具备超人级诊断能力。

Superhuman performance of a large language model on the reasoning tasks of a physician

  • 用医生专家评估大语言模型在五类临床推理任务中的表现。
  • 在全部实验中,模型表现均优于数百名执业医师,且持续优于前代AI。
  • 适合医疗决策支持、AI临床应用研究者关注,推动真实世界验证。

1959年Ledley和Lusted提出复杂临床诊断案例作为医学计算系统评估的黄金标准,至今仍被沿用。本文报告了大语言模型(LLM)在挑战性临床病例上的医生评估结果,对比基准为数百名执业医师。我们设计五项实验,涵盖鉴别诊断生成、推理展示、分诊排序、概率推理与管理推理,均由经验证的心理测量学医生专家裁定。此外,我们在波士顿一家大型三甲医学中心急诊科开展真实世界研究,比较人类专家与AI第二意见在三个预设诊断节点的表现:急诊分诊、初步评估及入院或入住重症监护。所有实验——包括病例题和急诊室第二意见——均显示该模型在诊断与推理能力上达到超人水平,且性能持续优于前代AI临床决策支持系统。研究结果表明,大语言模型已在通用医学诊断与管理推理任务中实现超人表现,实现了Ledley和Lusted提出的愿景,并迫切需要开展前瞻性试验。

原文摘要 · Abstract (English)

A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials.

医学AI大模型临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。