用多源医学数据提升视觉语言模型准确性,解决诊断报告错误问题
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
- 通过异构数据检索框架,精准匹配不同模态的医学报告
- 在11个数据集上显著提升模型事实准确性与输出可靠性
- 适合医疗AI研发者、临床辅助系统开发者使用
医学大视觉语言模型在临床应用中展现潜力,但存在事实性错误和不可靠输出,威胁真实诊疗安全。尽管检索增强生成(RAG)被视为潜在解决方案,现有医学多模态RAG系统难以在异构数据源间有效检索。检索结果与报告不相关会损害分析的真实性,知识不足则影响临床决策可信度。为此,我们构建了包含广泛多模态报告库与多样化文本语料的MedAtlas。基于此,提出HeteroRAG框架,通过模态特定的CLIP实现高效报告检索,并设计多语料查询生成器适配不同语料库。结合多源知识进行异构知识偏好微调,实现跨模态与多源知识对齐。在11个数据集、3种模态上的大量实验表明,HeteroRAG在多数医学视觉语言基准上达到当前最优性能,显著提升医学大模型的事实准确性和可靠性。
原文摘要 · Abstract (English)
Medical large vision-language Models (Med-LVLMs) have shown promise in clinical applications but suffer from factual inaccuracies and unreliable outputs, posing risks in real-world diagnostics. While RAG has emerged as a potential solution, current medical multimodal RAG systems are unable to perform effective retrieval across heterogeneous sources. The irrelevance of retrieved reports undermines the factuality of analysis, while insufficient knowledge affects the credibility of clinical decision-making. To bridge the research gap, we construct MedAtlas, which includes extensive multimodal report repositories and diverse text corpora. Based on it, we present HeteroRAG, a novel framework that enhances Med-LVLMs through heterogeneous knowledge sources. The framework introduces Modality-specific CLIPs for effective report retrieval and a Multi-corpora Query Generator for tailoring queries to diverse corpora. Incorporating knowledge from such multifaceted sources, Heterogeneous Knowledge Preference Tuning is performed to achieve cross-modality and multi-source knowledge alignment. Extensive experiments across 11 datasets and 3 modalities demonstrate that HeteroRAG achieves state-of-the-art performance in most medical vision language benchmarks, significantly improving factual accuracy and reliability of Med-LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。