用400个真实病例测试健康AI聊天机器人,发现其诊断准确率超80%。
A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI
- 用AI模拟患者与健康AI对话,批量测试诊断能力。
- 诊断准确率达81.8%(前1名),比传统工具少问47%问题。
- 适合关注医疗AI评估标准的研究者和开发者。
诊断错误在医疗中仍是个严峻挑战,越来越多患者依赖在线资源获取健康信息。尽管基于AI的医疗聊天机器人展现出潜力,但缺乏标准化、可扩展的评估框架。本研究提出一种可扩展的基准测试方法,用于评估健康AI系统,并通过一个名为August的对话式聊天机器人进行验证。该方法采用400个经验证的临床案例,覆盖14个医学专科,利用AI驱动的患者角色模拟真实临床交互。系统测试显示,August在400例中实现81.8%(327/400)的首诊准确率和85.0%(340/400)的前两名准确率,显著优于传统症状检查工具。系统在专科转诊上达到95.8%准确率,平均仅需16个问题(相较传统工具29个),减少47%提问量,同时保持共情对话。研究证明了聊天机器人提升医疗效率的潜力,但实际应用仍面临真实场景验证与客观数据整合等挑战。该方法为医疗AI评估提供了可复现的框架,有助于推动其在临床环境中的负责任发展。
原文摘要 · Abstract (English)
Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no standardized and scalable framework for evaluating their diagnostic capabilities. This study introduces a scalable benchmarking methodology for assessing health AI systems and demonstrates its application through August, an AI-driven conversational chatbot. Our methodology employs 400 validated clinical vignettes across 14 medical specialties, using AI-powered patient actors to simulate realistic clinical interactions. In systematic testing, August achieved a top-one diagnostic accuracy of 81.8% (327/400 cases) and a top-two accuracy of 85.0% (340/400 cases), significantly outperforming traditional symptom checkers. The system demonstrated 95.8% accuracy in specialist referrals and required 47% fewer questions compared to conventional symptom checkers (mean 16 vs 29 questions), while maintaining empathetic dialogue throughout consultations. These findings demonstrate the potential of AI chatbots to enhance healthcare delivery, though implementation challenges remain regarding real-world validation and integration of objective clinical data. This research provides a reproducible framework for evaluating healthcare AI systems, contributing to the responsible development and deployment of AI in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。