arXiv:2410.08431cs.CLcs.AI2024-10被引 2

用医学指南增强大模型,快速生成准确的术前指导。

oRetrieval Augmented Generation for 10 Large Language Models and its Generalizability in Assessing Medical Fitness

  • 用本地与国际指南构建知识库,通过检索增强生成术前建议。
  • GPT4-RAG准确率达96.4%,快20倍且无幻觉,媲美医生。
  • 适合医疗效率提升场景,尤其需快速标准化指导的手术准备。

大型语言模型在医疗应用中潜力巨大,但常缺乏专业临床知识。检索增强生成(RAG)可通过领域特定信息定制,适用于医疗场景。本研究评估了RAG模型在判断手术适宜性及提供术前指导中的准确性、一致性与安全性。基于35个本地和23个国际术前指南,构建了LLM-RAG模型,测试其对14个临床情景、7个术前指导维度的响应,共生成3,682条回答。采用Llamaindex处理临床文档,评估了包括GPT3.5、GPT4和Claude-3在内的10个大模型。以权威指南和专家判断为金标准,人类生成答案作为对比。所有LLM-RAG模型响应时间均在20秒内,显著快于临床医生(10分钟)。GPT4-RAG模型准确率达96.4%(人类为86.6%,p=0.016),无幻觉,生成内容与医生相当。结果在本地与国际指南间保持一致。研究证实,LLM-RAG在术前任务中具备高效、可扩展与可靠潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show potential for medical applications but often lack specialized clinical knowledge. Retrieval Augmented Generation (RAG) allows customization with domain-specific information, making it suitable for healthcare. This study evaluates the accuracy, consistency, and safety of RAG models in determining fitness for surgery and providing preoperative instructions. We developed LLM-RAG models using 35 local and 23 international preoperative guidelines and tested them against human-generated responses. A total of 3,682 responses were evaluated. Clinical documents were processed using Llamaindex, and 10 LLMs, including GPT3.5, GPT4, and Claude-3, were assessed. Fourteen clinical scenarios were analyzed, focusing on seven aspects of preoperative instructions. Established guidelines and expert judgment were used to determine correct responses, with human-generated answers serving as comparisons. The LLM-RAG models generated responses within 20 seconds, significantly faster than clinicians (10 minutes). The GPT4 LLM-RAG model achieved the highest accuracy (96.4% vs. 86.6%, p=0.016), with no hallucinations and producing correct instructions comparable to clinicians. Results were consistent across both local and international guidelines. This study demonstrates the potential of LLM-RAG models for preoperative healthcare tasks, highlighting their efficiency, scalability, and reliability.

医疗AIRAG大模型术前评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。