用大模型自动筛选手术患者是否需联合管理,准确率超90%。
Deployment and Evaluation of an EHR-integrated, Large Language Model-Powered Tool to Triage Surgical Patients
- 基于电子病历和临床标准,大模型自动判断患者是否适合联合管理。
- 敏感度达94%,特异性74%,误判多因流程或标准问题而非模型错误。
- 适合想自动化临床筛查流程的医院和医疗AI研究者。
外科联合管理(SCM)是一种基于证据的模式,由住院医生与外科团队共同管理术前术后病情复杂的患者。尽管具有临床和经济价值,但受限于需手动识别符合条件的患者。为验证能否自动化筛查,我们在斯坦福医疗中心开展前瞻性、非盲法研究,部署了一款基于大语言模型、集成电子病历的筛查工具(SCM Navigator),提供SCM建议并由医生复核。该工具利用术前记录、结构化数据及围术期并发症临床标准,将患者分为合适、不合适或可能合适三类。主治医师给出临床判断并提供自由文本反馈。以医生判断为金标准,计算敏感度、特异度、阳性预测值和阴性预测值。对不一致案例进行主题分析,并对所有假阴性病例及30个假阳性最多类别样本进行人工病历审查。自上线以来,共筛查6,193例,其中1,582例(23%)被推荐会诊。结果显示,系统敏感度为0.94(95%置信区间0.91–0.96),特异度为0.74(95%置信区间0.71–0.77)。事后病历审查表明,多数差异源于可改进的临床标准、机构流程或医生实践差异,而非大模型误判,仅2/19(11%)假阴性由模型错误导致。结果表明,一种集成大模型、嵌入电子病历、人工参与闭环的AI系统可安全、准确地对外科患者进行联合管理筛查,且有望辅助甚至实现耗时临床流程的自动化。
原文摘要 · Abstract (English)
Surgical co-management (SCM) is an evidence-based model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite its clinical and financial value, SCM is limited by the need to manually identify eligible patients. To determine whether SCM triage can be automated, we conducted a prospective, unblinded study at Stanford Health Care in which an LLM-based, electronic health record (EHR)-integrated triage tool (SCM Navigator) provided SCM recommendations followed by physician review. Using pre-operative documentation, structured data, and clinical criteria for perioperative morbidity, SCM Navigator categorized patients as appropriate, not appropriate, or possibly appropriate for SCM. Faculty indicated their clinical judgment and provided free-text feedback when they disagreed. Sensitivity, specificity, positive predictive value, and negative predictive value were measured using physician determinations as a reference. Free-text reasons were thematically categorized, and manual chart review was conducted on all false-negative cases and 30 randomly selected cases from the largest false-positive category. Since deployment, 6,193 cases have been triaged, of which 1,582 (23%) were recommended for hospitalist consultation. SCM Navigator displayed high sensitivity (0.94, 95% CI 0.91-0.96) and moderate specificity (0.74, 95% CI 0.71-0.77). Post-hoc chart review suggested most discrepancies reflect modifiable gaps in clinical criteria, institutional workflow, or physician practice variability rather than LLM misclassification, which accounted for 2 of 19 (11%) false-negative cases. These findings demonstrate that an LLM-powered, EHR-integrated, human-in-the-loop AI system can accurately and safely triage surgical patients for SCM, and that AI-enabled screening tools can augment and potentially automate time-intensive clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。