arXiv:2411.09213cs.CLcs.AI2024-11被引 17

构建医疗RAG评估基准,测试模型在真实场景下的可靠性。

Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering

  • 提出MedRGB基准,涵盖检索不足、信息融合等实用场景
  • 实测主流模型在噪声文档下准确率下降超40%
  • 适合医疗AI研发者与评测人员参考

检索增强生成(RAG)已成为提升大语言模型(LLMs)在医学等知识密集型任务中表现的有前景方法。然而,医学领域的敏感性要求系统必须完全准确可信。现有RAG基准主要聚焦标准的检索-回答流程,忽视了衡量可靠医疗系统的关键实际场景。本文填补这一空白,提出一个全面的医疗问答(QA)系统RAG评估框架,涵盖充分性、集成性和鲁棒性等维度。我们构建了医学检索增强生成基准(MedRGB),为四个医学QA数据集添加多种补充元素,用于测试LLMs在特定场景下的能力。利用MedRGB,我们在多种检索条件下对当前主流商业模型和开源模型进行了广泛评估。实验结果揭示了现有模型在处理检索文档中的噪声和错误信息时能力有限。我们进一步分析了LLMs的推理过程,提供了有价值的洞察与未来发展方向。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) has emerged as a promising approach to enhance the performance of large language models (LLMs) in knowledge-intensive tasks such as those from medical domain. However, the sensitive nature of the medical domain necessitates a completely accurate and trustworthy system. While existing RAG benchmarks primarily focus on the standard retrieve-answer setting, they overlook many practical scenarios that measure crucial aspects of a reliable medical system. This paper addresses this gap by providing a comprehensive evaluation framework for medical question-answering (QA) systems in a RAG setting for these situations, including sufficiency, integration, and robustness. We introduce Medical Retrieval-Augmented Generation Benchmark (MedRGB) that provides various supplementary elements to four medical QA datasets for testing LLMs' ability to handle these specific scenarios. Utilizing MedRGB, we conduct extensive evaluations of both state-of-the-art commercial LLMs and open-source models across multiple retrieval conditions. Our experimental results reveals current models' limited ability to handle noise and misinformation in the retrieved documents. We further analyze the LLMs' reasoning processes to provides valuable insights and future directions for developing RAG systems in this critical medical domain.

医疗AIRAG评估大模型知识检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。