用医学证据异质性分析法,检测医疗RAG系统回答中的错误
M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
- 基于循证医学的异质性分析,从多源文献验证RAG回答一致性
- 在多个LLM上提升23.31%的准确率,有效识别幻觉与知识误用
- 适合医疗AI验证、临床辅助决策系统开发者使用
检索增强生成(RAG)通过整合大语言模型(LLMs)与外部医学文献,在提升医疗问答系统表现方面展现出潜力。然而,现有RAG应用仍存在生成错误信息(如幻觉)和未能正确利用外部知识的问题。为此,本文提出M-Eval方法,受循证医学中异质性分析的启发,通过多源文献证据检验RAG输出的一致性与事实准确性。首先从外部知识库提取额外医学文献,再检索RAG生成的证据文档,利用异质性分析判断这些证据是否支持回答中的不同观点。该方法不仅能验证回答准确性,还可评估证据可靠性。实验表明,该方法在多种LLM上将准确率最高提升23.31%,有助于发现当前医疗RAG系统的错误,提升大模型在医疗场景下的可信度与诊断安全性。
原文摘要 · Abstract (English)
Retrieval-augmented Generation (RAG) has demonstrated potential in enhancing medical question-answering systems through the integration of large language models (LLMs) with external medical literature. LLMs can retrieve relevant medical articles to generate more professional responses efficiently. However, current RAG applications still face problems. They generate incorrect information, such as hallucinations, and they fail to use external knowledge correctly. To solve these issues, we propose a new method named M-Eval. This method is inspired by the heterogeneity analysis approach used in Evidence-Based Medicine (EBM). Our approach can check for factual errors in RAG responses using evidence from multiple sources. First, we extract additional medical literature from external knowledge bases. Then, we retrieve the evidence documents generated by the RAG system. We use heterogeneity analysis to check whether the evidence supports different viewpoints in the response. In addition to verifying the accuracy of the response, we also assess the reliability of the evidence provided by the RAG system. Our method shows an improvement of up to 23.31% accuracy across various LLMs. This work can help detect errors in current RAG-based medical systems. It also makes the applications of LLMs more reliable and reduces diagnostic errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。