arXiv:2509.12194cs.AIcs.CV2025-09被引 2

AI医生能像专家一样做诊断推理,且被误认为是真人。

Teaching large language models to reason like expert diagnosticians

  • 用幻灯片形式模拟专家诊断过程,从病历自动生成推理报告。
  • 在72例疑难病例中,69%仅凭病历就准确找出病因。
  • 首次获《新英格兰医学杂志》刊登,适合医疗AI研究者参考。

鉴别诊断是一个整合患者信息与医学知识的迭代过程。自1923年以来持续发布的《新英格兰医学杂志临床病理研讨会》(NEJM CPCs)包含专家医生向同行展示诊断推理的案例,长期用于评估AI系统。然而以往评估多关注最终诊断准确率,忽视诊断推理的精细过程。本文提出Dr. CaBot,一个代理型AI系统,仅凭初始病例描述即可生成文字与语音驱动的幻灯片式诊断报告。该系统成为首个在百年历史的NEJM CPCs中发表的AI诊断。盲评中,46/62(74%)次试验中医生无法分辨报告来源(AI或人类),且在质量维度上给予正面评价。在处理美国国立卫生研究院未确诊疾病网络(NIH Undiagnosed Diseases Network)72例疑难病例时,仅基于转诊记录即成功识别出50/72(69%)例工作诊断。为促进透明与研究,我们构建了CPC-Bench,一个基于7,102个CPC案例、47,648道题目、涵盖10项任务的医师验证基准。结果显示CaBot优于前沿模型,相关代码与数据集已公开。

原文摘要 · Abstract (English)

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert physicians who demonstrate diagnostic reasoning to peers, and have been used for decades to evaluate AI. However, prior AI evaluations have largely focused on final diagnostic accuracy rather than nuanced clinical reasoning. Here, we introduce Dr. CaBot, an agentic AI system that emulates an expert diagnostician by generating written and narrated slide-based presentations from an initial case description alone. CaBot recently generated the first AI diagnosis published in the 100+ year history of the NEJM CPCs. In blinded evaluations, physicians misclassified the source of the differential (CaBot vs. physician-written) in 46/62 (74%) of trials and rated them favorably across quality dimensions. When tasked with solving cases for 72 patients with undiagnosed disease from the NIH Undiagnosed Diseases Network, CaBot identified the working diagnosis in 50/72 (69%) of cases from referral notes alone. To promote transparency and research, we also developed CPC-Bench, a physician-validated benchmark based on 7,102 CPCs and 47,648 questions across 10 tasks. We show that CaBot outperforms frontier models on CPC-Bench, and release both CaBot and CPC-Bench publicly to foster progress in clinical AI.

医疗AI诊断推理大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。