用大模型精准提取医院指南信息,回答临床问题又快又准。
Clean & Clear: Feasibility of Safe LLM Clinical Guidance
- 基于Llama-3.1-8B模型,直接引用指南原文,不靠推理生成答案。
- 97%的问答相关性达标,召回率达100%,关键信息零遗漏。
- 医生评估显示72%回复无逻辑错误,响应速度比人快20秒。
背景:临床指南是现代医疗中安全循证医学的核心,涵盖诊断标准、治疗方案与监测建议。大型语言模型驱动的聊天机器人在医疗问答任务中展现出巨大潜力,可快速准确回应医学咨询。本研究目标是开发并初步评估一款基于伦敦大学学院医院(UCLH)临床指南的LLM聊天机器人,实现对临床指南问题的可靠回答。方法:采用开源权重的Llama-3.1-8B模型,从指南中提取相关信息以回答问题,强调引用原始信息而非生成解释。由七名病房医生通过与标准答案对比评估聊天机器人的表现。结果:聊天机器人在相关性方面表现良好,约73%的回答被评为高度相关,显示出对临床情境的良好理解;其提取指南条目的召回率达到1.00,显著降低遗漏关键信息的风险;约78%的回答在完整性上被评定为满意;仅有约14.5%的回答包含少量冗余信息,表明精度略有不足。平均响应时间为10秒,远低于人类平均30秒。临床推理评估显示,72%的回答无逻辑缺陷。该聊天机器人在加速获取本地化临床信息方面具有显著潜力。
原文摘要 · Abstract (English)
Background: Clinical guidelines are central to safe evidence-based medicine in modern healthcare, providing diagnostic criteria, treatment options and monitoring advice for a wide range of illnesses. LLM-empowered chatbots have shown great promise in Healthcare Q&A tasks, offering the potential to provide quick and accurate responses to medical inquiries. Our main objective was the development and preliminary assessment of an LLM-empowered chatbot software capable of reliably answering clinical guideline questions using University College London Hospital (UCLH) clinical guidelines. Methods: We used the open-weight Llama-3.1-8B LLM to extract relevant information from the UCLH guidelines to answer questions. Our approach highlights the safety and reliability of referencing information over its interpretation and response generation. Seven doctors from the ward assessed the chatbot's performance by comparing its answers to the gold standard. Results: Our chatbot demonstrates promising performance in terms of relevance, with ~73% of its responses rated as very relevant, showcasing a strong understanding of the clinical context. Importantly, our chatbot achieves a recall of 1.00 for extracted guideline lines, substantially minimising the risk of missing critical information. Approximately 78% of responses were rated satisfactory in terms of completeness. A small portion (~14.5%) contained minor unnecessary information, indicating occasional lapses in precision. The chatbot' showed high efficiency, with an average completion time of 10 seconds, compared to 30 seconds for human respondents. Evaluation of clinical reasoning showed that 72% of the chatbot's responses were without flaws. Our chatbot demonstrates significant potential to speed up and improve the process of accessing locally relevant clinical information for healthcare professionals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。