arXiv:2603.08448cs.HCcs.AI2026-03被引 8

AI助手在门诊提前问诊,90%匹配最终诊断,安全可用。

A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

  • AI通过聊天收集病史并生成诊断建议,实时由人工监督。
  • 90%病例包含最终诊断,75%在前3个推荐中,与医生水平相当。
  • 患者和医生都认可其效果,尤其提升就诊准备度,适合临床落地。

大型语言模型(LLM)驱动的对话式AI在模拟环境中展现了面向患者的诊断与管理能力。本研究开展一项前瞻性单臂可行性研究,评估基于LLM的对话式AI系统Articulate Medical Intelligence Explorer(AMIE)在一家顶尖医学中心的急症门诊中的实际应用。100名成年患者在就诊前最多5天内完成与AMIE的文字对话。研究评估了对话安全性、质量、患者与医生体验及临床推理能力,对比初级保健医生(PCPs)。所有患者-AMIE交互均由人工安全监督员实时监控,未因预设标准而中断任何会话。患者满意度高,对AI态度显著改善(p < 0.001)。PCPs认为AMIE输出有用,并提升了就诊准备度。病历回顾8周后显示,AMIE的鉴别诊断(DDx)包含最终诊断的比例为90%,前3项准确率达75%。盲评显示,AMIE与PCP的DDx和管理计划(Mx)整体质量无显著差异(DDx p = 0.6;Mx恰当性 p = 0.1,安全性 p = 1.0)。但医生在管理方案的实用性(p = 0.003)和成本效益(p = 0.004)上表现更优。尽管需进一步研究,该研究证实了对话式AI在真实场景中的可行性、安全性与用户接受度,是迈向临床转化的重要一步。

原文摘要 · Abstract (English)

Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clinical practice requires assessment in real-world workflows with rigorous safety oversight. We report a prospective, single-arm feasibility study of an LLM-based conversational AI, the Articulate Medical Intelligence Explorer (AMIE), conducting clinical history taking and presentation of potential diagnoses for patients to discuss with their provider at urgent care appointments at a leading academic medical center. 100 adult patients completed an AMIE text-chat interaction up to 5 days before their appointment. We sought to assess the conversational safety and quality, patient and clinician experience, and clinical reasoning capabilities compared to primary care providers (PCPs). Human safety supervisors monitored all patient-AMIE interactions in real time and did not need to intervene to stop any consultations based on pre-defined criteria. Patients reported high satisfaction and their attitudes towards AI improved after interacting with AMIE (p < 0.001). PCPs found AMIE's output useful with a positive impact on preparedness. AMIE's differential diagnosis (DDx) included the final diagnosis, per chart review 8 weeks post-encounter, in 90% of cases, with 75% top-3 accuracy. Blinded assessment of AMIE and PCP DDx and management (Mx) plans suggested similar overall DDx and Mx plan quality, without significant differences for DDx (p = 0.6) and appropriateness and safety of Mx (p = 0.1 and 1.0, respectively). PCPs outperformed AMIE in the practicality (p = 0.003) and cost effectiveness (p = 0.004) of Mx. While further research is needed, this study demonstrates the initial feasibility, safety, and user acceptance of conversational AI in a real-world setting, representing crucial steps towards clinical translation.

对话AI临床辅助大模型医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。