医生监督的AI聊天助手在真实医疗场景中提升患者体验,安全可靠。
Conversational Medical AI: Ready for Practice
- 医生监督下运行的LLM对话系统接入真实医疗咨询平台。
- 926例随机实验中,患者满意度和信息清晰度显著提升,信任感不变。
- 81%患者愿意使用,95%对话获医生评为优质,适合临床落地。
医生短缺导致医疗资源获取困难。尽管对话式人工智能有望缓解这一问题,但其在真实医疗场景中面向患者的安全部署仍缺乏研究。本文首次在真实医疗环境中对基于大语言模型的医生监督型对话代理进行大规模评估。该代理“Mo”被集成至现有医疗咨询服务中。在为期三周的随机对照实验中,共处理926例案例,其中298例完整交互由医生评估安全性和医学准确性。患者报告信息清晰度(3.73 vs 3.62,满分4,p<0.05)与整体满意度(4.58 vs 4.42,满分5,p<0.05)均高于标准护理,且信任感与共情感知无显著差异。81%受访者选择使用,超过以往医疗AI接受度基准。医生监督确保安全性,95%对话获全科医生评定为“良好”或“优秀”。结果表明,经精心设计的AI医疗助手可在保障安全前提下改善患者体验,并为现有医疗体系整合提供实证支持。
原文摘要 · Abstract (English)
The shortage of doctors is creating a critical squeeze in access to medical expertise. While conversational Artificial Intelligence (AI) holds promise in addressing this problem, its safe deployment in patient-facing roles remains largely unexplored in real-world medical settings. We present the first large-scale evaluation of a physician-supervised LLM-based conversational agent in a real-world medical setting. Our agent, Mo, was integrated into an existing medical advice chat service. Over a three-week period, we conducted a randomized controlled experiment with 926 cases to evaluate patient experience and satisfaction. Among these, Mo handled 298 complete patient interactions, for which we report physician-assessed measures of safety and medical accuracy. Patients reported higher clarity of information (3.73 vs 3.62 out of 4, p < 0.05) and overall satisfaction (4.58 vs 4.42 out of 5, p < 0.05) with AI-assisted conversations compared to standard care, while showing equivalent levels of trust and perceived empathy. The high opt-in rate (81% among respondents) exceeded previous benchmarks for AI acceptance in healthcare. Physician oversight ensured safety, with 95% of conversations rated as "good" or "excellent" by general practitioners experienced in operating a medical advice chat service. Our findings demonstrate that carefully implemented AI medical assistants can enhance patient experience while maintaining safety standards through physician supervision. This work provides empirical evidence for the feasibility of AI deployment in healthcare communication and insights into the requirements for successful integration into existing healthcare services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。