arXiv:2412.18096cs.AI2024-12被引 1

PEACH是用于术前决策的AI聊天机器人,实测准确率达97.5%。

Real-world Deployment and Evaluation of PErioperative AI CHatbot (PEACH) -- a Large Language Model Chatbot for Perioperative Medicine

  • 将35项院内术前指南嵌入安全LLM框架,实现临床决策支持
  • 240次真实使用中准确率96.7%,幻觉与偏差极低
  • 医生称其显著提速决策,适合临床高效辅助场景

大型语言模型(LLMs)在医疗领域日益成为复杂专业任务的强大工具。本研究介绍了术前AI聊天机器人PEACH的开发与评估,该系统基于安全的Claude 3.5 Sonet LLM框架(由新加坡政府开发的Pair Chat平台),集成35项机构术前指南,用于支持术前临床决策。PEACH在真实世界数据下进行了无声部署测试,评估了准确性、安全性与可用性。偏离与幻觉按潜在危害分类,用户反馈通过技术接受模型(TAM)评估。初始部署后对一项协议进行更新。在240次真实临床交互中,PEACH首次生成准确率为97.5%(78/80),整体准确率为96.7%(232/240),三次迭代后更新版本准确率达97.9%(235/240),显著高于95%零假设(p = 0.018,95% CI: 0.952-0.991)。幻觉与偏离均极低(分别为1/240和2/240)。临床医生报告95%情况下决策速度加快,组内一致性κ值为0.772–0.893,主治医师间为0.610–0.784。PEACH是精准、可适应的工具,能提升术前决策的一致性与效率。未来研究应探索其跨专科扩展及对临床结局的影响。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are emerging as powerful tools in healthcare, particularly for complex, domain-specific tasks. This study describes the development and evaluation of the PErioperative AI CHatbot (PEACH), a secure LLM-based system integrated with local perioperative guidelines to support preoperative clinical decision-making. PEACH was embedded with 35 institutional perioperative protocols in the secure Claude 3.5 Sonet LLM framework within Pair Chat (developed by Singapore Government) and tested in a silent deployment with real-world data. Accuracy, safety, and usability were assessed. Deviations and hallucinations were categorized based on potential harm, and user feedback was evaluated using the Technology Acceptance Model (TAM). Updates were made after the initial silent deployment to amend one protocol. In 240 real-world clinical iterations, PEACH achieved a first-generation accuracy of 97.5% (78/80) and an overall accuracy of 96.7% (232/240) across three iterations. The updated PEACH demonstrated improved accuracy of 97.9% (235/240), with a statistically significant difference from the null hypothesis of 95% accuracy (p = 0.018, 95% CI: 0.952-0.991). Minimal hallucinations and deviations were observed (both 1/240 and 2/240, respectively). Clinicians reported that PEACH expedited decisions in 95% of cases, and inter-rater reliability ranged from kappa 0.772-0.893 within PEACH and 0.610-0.784 among attendings. PEACH is an accurate, adaptable tool that enhances consistency and efficiency in perioperative decision-making. Future research should explore its scalability across specialties and its impact on clinical outcomes.

AI医疗临床决策大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。