arXiv:2501.02460cs.CL2025-01ACL被引 25

解决医学大模型知识幻觉,通过智能规划多源知识检索

Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications

  • 将多源医学知识检索建模为源规划问题,精准匹配不同知识源特性
  • 在多个医学数据集上显著提升多源知识利用效率,超越现有方法
  • 适合医疗AI研发者、临床决策系统开发者使用

大型语言模型在医学诊断推理、科研知识获取、临床决策和公众健康咨询支持方面具有潜力,但受限于医学知识不足,常产生幻觉。引入外部知识至关重要,需实现多源知识获取。本文将该问题建模为源规划问题,即根据上下文生成适配不同知识源特性的查询。现有方法或忽略源规划,或因模型预期与实际内容不匹配而效果不佳。为此,我们构建了MedOmniKB——一个涵盖多种类型和结构的医学知识源库,并提出源规划优化方法,通过专家模型探索评估潜在规划方案,训练小模型学习源对齐。实验表明,该方法显著提升多源规划性能,使优化后的小模型在利用多样化医学知识源方面达到当前最优水平。

原文摘要 · Abstract (English)

Large language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer health inquiry support. However, they often generate hallucinations due to limited medical knowledge. Incorporating external knowledge is therefore critical, which necessitates multi-source knowledge acquisition. We address this challenge by framing it as a source planning problem, which is to formulate context-appropriate queries tailored to the attributes of diverse sources. Existing approaches either overlook source planning or fail to achieve it effectively due to misalignment between the model's expectation of the sources and their actual content. To bridge this gap, we present MedOmniKB, a repository comprising multigenre and multi-structured medical knowledge sources. Leveraging these sources, we propose the Source Planning Optimisation method, which enhances multi-source utilisation. Our approach involves enabling an expert model to explore and evaluate potential plans while training a smaller model to learn source alignment. Experimental results demonstrate that our method substantially improves multi-source planning performance, enabling the optimised small model to achieve state-of-the-art results in leveraging diverse medical knowledge sources.

医学AI知识增强多源检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。