AI医生在急诊场景中表现堪比真人医生,诊断准确率超八成。
Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting
- 用多智能体大模型构建自主诊疗AI,模拟真实急诊问诊流程。
- 诊断匹配率达81%,治疗方案一致率99.2%,无虚假诊断。
- 部分病例表现优于人类医生,适合缓解医疗人力短缺问题。
全球预计到2030年将面临1100万医疗从业者缺口,且临床工作50%时间耗费于行政事务。人工智能有望缓解此困境,但迄今尚无端到端的自主大语言模型(LLM)系统在真实临床环境中被严格评估。本研究回顾性比较了多智能体LLM框架Doctronic与执业医师在500例连续远程急诊就诊中的表现。主要终点包括诊断一致性、治疗方案一致性及安全指标,由盲法大模型判别与专家人工评审共同评估。结果显示,Doctronic与医生的首要诊断匹配率为81%,治疗方案一致率达99.2%,未出现临床幻觉(如无依据的诊断或治疗)。在分歧案例的专家评审中,AI表现更优占36.1%,人类更优占9.3%,其余情况相当。结论表明,该自主AI医生在大规模真实场景中达到与人类医生相当甚至部分超越的临床决策水平,为应对医疗人力短缺提供可行方案。
原文摘要 · Abstract (English)
Background: Globally we face a projected shortage of 11 million healthcare practitioners by 2030, and administrative burden consumes 50% of clinical time. Artificial intelligence (AI) has the potential to help alleviate these problems. However, no end-to-end autonomous large language model (LLM)-based AI system has been rigorously evaluated in real-world clinical practice. In this study, we evaluated whether a multi-agent LLM-based AI framework can function autonomously as an AI doctor in a virtual urgent care setting. Methods: We retrospectively compared the performance of the multi-agent AI system Doctronic and board-certified clinicians across 500 consecutive urgent-care telehealth encounters. The primary end points: diagnostic concordance, treatment plan consistency, and safety metrics, were assessed by blinded LLM-based adjudication and expert human review. Results: The top diagnosis of Doctronic and clinician matched in 81% of cases, and the treatment plan aligned in 99.2% of cases. No clinical hallucinations occurred (e.g., diagnosis or treatment not supported by clinical findings). In an expert review of discordant cases, AI performance was superior in 36.1%, and human performance was superior in 9.3%; the diagnoses were equivalent in the remaining cases. Conclusions: In this first large-scale validation of an autonomous AI doctor, we demonstrated strong diagnostic and treatment plan concordance with human clinicians, with AI performance matching and in some cases exceeding that of practicing clinicians. These findings indicate that multi-agent AI systems achieve comparable clinical decision-making to human providers and offer a potential solution to healthcare workforce shortages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。