用多模态检索增强生成技术,提升阿尔茨海默病临床决策支持
AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles
- 融合文本与图像的跨模态注意力机制,提升医学文献检索效率
- 在PubMedQA等数据集上表现优于基准模型,生成结果与人类水平相当
- 适合临床医生快速获取精准医学证据,减少AI幻觉
近年来,生成式AI推动了大型语言模型(LLMs)的发展,使其能够整合多种数据类型以支持决策。其中,多模态检索增强生成(RAG)应用前景广阔,结合信息检索与生成模型优势,在临床等场景中具有高实用性。本文提出AlzheimerRAG,一种面向阿尔茨海默病临床案例的多模态RAG系统,基于PubMed文章构建。该系统采用跨模态注意力融合技术,高效索引并访问海量生物医学文献,实现文本与视觉信息的协同处理。实验结果显示,其在BioASQ和PubMedQA等基准测试中表现优于现有方法,能准确检索并合成领域特定信息。此外,通过多个阿尔茨海默病临床场景案例研究,验证了AlzheimerRAG生成内容与人类水平相当,且幻觉率低。
原文摘要 · Abstract (English)
Recent advancements in generative AI have fostered the development of highly adept Large Language Models (LLMs) that integrate diverse data types to empower decision-making. Among these, multimodal retrieval-augmented generation (RAG) applications are promising because they combine the strengths of information retrieval and generative models, enhancing their utility across various domains, including clinical use cases. This paper introduces AlzheimerRAG, a Multimodal RAG application for clinical use cases, primarily focusing on Alzheimer's Disease case studies from PubMed articles. This application incorporates cross-modal attention fusion techniques to integrate textual and visual data processing by efficiently indexing and accessing vast amounts of biomedical literature. Our experimental results, compared to benchmarks such as BioASQ and PubMedQA, have yielded improved performance in the retrieval and synthesis of domain-specific information. We also present a case study using our multimodal RAG in various Alzheimer's clinical scenarios. We infer that AlzheimerRAG can generate responses with accuracy non-inferior to humans and with low rates of hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。