医学领域用检索增强生成技术,提升大模型临床应用能力。
Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- 结合外部知识检索与生成,弥补通用大模型医学知识不足。
- 现有研究多用公开数据,医疗专用模型和评估仍较薄弱。
- 适合关注医疗AI落地、伦理安全与跨语言适配的研究者。
医学知识快速积累与临床实践日益复杂,给诊疗带来挑战。尽管大型语言模型(LLMs)展现出潜力,但其内在局限依然存在。检索增强生成(RAG)技术有望提升其在临床中的实用性。本研究综述了医学领域RAG的技术实现、临床应用及伦理考量。发现当前研究主要依赖公开数据,私有数据应用有限;检索环节多采用以英语为中心的嵌入模型,而生成模型普遍为通用型,医疗专用模型使用较少;评估方面,自动化指标关注生成质量与任务表现,人工评估侧重准确性、完整性、相关性与流畅性,但对偏见与安全性关注不足。RAG应用集中于问答、报告生成、文本摘要与信息抽取。总体而言,医学RAG尚处早期阶段,亟需在临床验证、跨语言适应及低资源场景支持方面取得进展,以实现可信赖、负责任的全球应用。
原文摘要 · Abstract (English)
The rapid growth of medical knowledge and increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) have demonstrated value; however, inherent limitations remain. Retrieval-augmented generation (RAG) technologies show potential to enhance their clinical applicability. This study reviewed RAG applications in medicine. We found that research primarily relied on publicly available data, with limited application in private data. For retrieval, approaches commonly relied on English-centric embedding models, while LLMs were mostly generic, with limited use of medical-specific LLMs. For evaluation, automated metrics evaluated generation quality and task performance, whereas human evaluation focused on accuracy, completeness, relevance, and fluency, with insufficient attention to bias and safety. RAG applications were concentrated on question answering, report generation, text summarization, and information extraction. Overall, medical RAG remains at an early stage, requiring advances in clinical validation, cross-linguistic adaptation, and support for low-resource settings to enable trustworthy and responsible global use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。