PULSE AI辅助医生诊断,对常见病罕见病表现稳定。
Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
- 结合医学大模型与文献检索,动态生成诊断推理
- 在82个真实病例中达到专家级准确率,罕见病不降性能
- 适合临床协作场景,但需警惕自动化偏见
我们提出PULSE,一个融合领域微调的大语言模型与科学文献检索的医疗推理代理,用于支持复杂真实病例的诊断决策。为评估其能力,我们构建了一个包含82例真实内分泌科病例报告的基准数据集,涵盖多种疾病类型和发病率水平。在受控实验中,我们将PULSE的表现与不同经验层次的医生(从住院医师到资深专家)进行对比,并分析AI辅助对人类诊断推理的影响。PULSE在Top@1和Top@4阈值下均达到专家级准确率,优于住院医师和初级专科医生,与资深专家相当;且在疾病罕见性增加时,其性能保持稳定,而人类医生的准确率则下降。该代理展现出自适应推理能力,输出长度随病例难度增加,类似专家的深度思考过程。协同使用时,PULSE帮助医生纠正初始错误并拓展诊断假设,但也存在自动化偏见风险。研究探讨了串行与并行协作流程,表明PULSE在常见与罕见病症中均能提供可靠支持。这些发现揭示了基于语言模型的代理在临床诊断中的潜力与局限,并为其实现真实世界决策支持提供了评估框架。
原文摘要 · Abstract (English)
We present PULSE, a medical reasoning agent that combines a domain-tuned large language model with scientific literature retrieval to support diagnostic decision-making in complex real-world cases. To evaluate its capabilities, we curated a benchmark of 82 authentic endocrinology case reports encompassing a broad spectrum of disease types and incidence levels. In controlled experiments, we compared PULSE's performance against physicians with varying levels of expertise-from residents to senior specialists-and examined how AI assistance influenced human diagnostic reasoning. PULSE attained expert-competitive accuracy, outperforming residents and junior specialists while matching senior specialist performance at both Top@1 and Top@4 thresholds. Unlike physicians, whose accuracy declined with disease rarity, PULSE maintained stable performance across incidence tiers. The agent also exhibited adaptive reasoning, increasing output length with case difficulty in a manner analogous to the longer deliberation observed among expert clinicians. When used collaboratively, PULSE enabled physicians to correct initial errors and broaden diagnostic hypotheses, but also introduced risks of automation bias. The study explores both serial and concurrent collaboration workflows, revealing that PULSE offers robust support across common and rare presentations. These findings underscore both the promise and the limitations of language model-based agents in clinical diagnosis, and offer a framework for evaluating their role in real-world decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。