AI医生在常见内科病诊断上准确率超医生,且更快更省钱。
Comparisons between a Large Language Model-based Real-Time Compound Diagnostic Medical AI Interface and Physicians for Common Internal Medicine Cases using Simulated Patients
- 用大模型构建实时复合诊断AI接口,模拟真实临床场景测试。
- AI首诊准确率达80%,两次诊断全对,比医生快44.6%、成本降98.1%。
- 适合想提升诊疗效率的基层医疗或初诊场景使用。
目的:开发基于大语言模型(LLM)的实时复合诊断医疗AI接口,并通过模拟患者在美国执业医师资格考试(USMLE)Step 2临床技能(CS)风格试题下的临床试验,比较该接口与医生在常见内科病例中的表现。方法:2024年8月20日开展非随机对照临床试验,招募1名全科医生、2名内科住院医师(二、三年级),以及5名模拟患者。临床案例改编自USMLE Step 2 CS题型,共设计10个基于真实患者的典型内科病例,包含初始诊断评估信息。主要结局为首次鉴别诊断的准确性,重复性通过一致率评估。结果:医生首次鉴别诊断准确率为50%~70%,而实时复合诊断AI接口准确率达80%。首次诊断一致性比例为0.7。医生首次及第二次诊断准确率为70%~90%,而AI接口达100%。AI平均耗时557秒,比医生平均1006秒减少44.6%;成本为0.08美元,相比医生均值4.2美元降低98.1%。医生护理满意度为4.2~4.3分,AI为3.9分。结论:基于大语言模型的实时复合诊断医疗AI接口在诊断准确性和患者满意度方面与医生相当,同时耗时更短、成本更低。研究提示此类AI接口可能有助于辅助常见内科病例的初级诊疗。
原文摘要 · Abstract (English)
Objective To develop an LLM based realtime compound diagnostic medical AI interface and performed a clinical trial comparing this interface and physicians for common internal medicine cases based on the United States Medical License Exam (USMLE) Step 2 Clinical Skill (CS) style exams. Methods A nonrandomized clinical trial was conducted on August 20, 2024. We recruited one general physician, two internal medicine residents (2nd and 3rd year), and five simulated patients. The clinical vignettes were adapted from the USMLE Step 2 CS style exams. We developed 10 representative internal medicine cases based on actual patients and included information available on initial diagnostic evaluation. Primary outcome was the accuracy of the first differential diagnosis. Repeatability was evaluated based on the proportion of agreement. Results The accuracy of the physicians' first differential diagnosis ranged from 50% to 70%, whereas the realtime compound diagnostic medical AI interface achieved an accuracy of 80%. The proportion of agreement for the first differential diagnosis was 0.7. The accuracy of the first and second differential diagnoses ranged from 70% to 90% for physicians, whereas the AI interface achieved an accuracy rate of 100%. The average time for the AI interface (557 sec) was 44.6% shorter than that of the physicians (1006 sec). The AI interface ($0.08) also reduced costs by 98.1% compared to the physicians' average ($4.2). Patient satisfaction scores ranged from 4.2 to 4.3 for care by physicians and were 3.9 for the AI interface Conclusion An LLM based realtime compound diagnostic medical AI interface demonstrated diagnostic accuracy and patient satisfaction comparable to those of a physician, while requiring less time and lower costs. These findings suggest that AI interfaces may have the potential to assist primary care consultations for common internal medicine cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。