构建医学领域RAG评估基准,测试大模型问答可靠性
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
- 设计跨语言医学RAG评测框架,整合维基和PubMed数据
- RAG显著提升医学问答可靠性,但长问题生成略降低可读性
- 开源工具包支持组件对比,适合医疗AI研究与开发
尽管检索增强生成(RAG)在科学与临床问答系统中迅速应用,但医学领域的全面评估基准仍显不足。为此,我们提出医学检索增强生成(MRAG)基准,涵盖英汉双语多种任务,并构建基于维基百科和PubMed的语料库。同时开发了MRAG-Toolkit,支持对不同RAG组件的系统性探索。实验表明:(a) RAG在各类MRAG任务中提升了大模型的可靠性;(b) RAG性能受检索方法、模型规模及提示策略影响;(c) 尽管提升了有用性和推理质量,长问题回答的可读性略有下降。论文接受后将按CCBY-4.0许可发布MRAG-Bench数据集与工具包,推动学术与产业应用。
原文摘要 · Abstract (English)
While Retrieval-Augmented Generation (RAG) has been swiftly adopted in scientific and clinical QA systems, a comprehensive evaluation benchmark in the medical domain is lacking. To address this gap, we introduce the Medical Retrieval-Augmented Generation (MRAG) benchmark, covering various tasks in English and Chinese languages, and building a corpus with Wikipedia and Pubmed. Additionally, we develop the MRAG-Toolkit, facilitating systematic exploration of different RAG components. Our experiments reveal that: (a) RAG enhances LLM reliability across MRAG tasks. (b) the performance of RAG systems is influenced by retrieval approaches, model sizes, and prompting strategies. (c) While RAG improves usefulness and reasoning quality, LLM responses may become slightly less readable for long-form questions. We will release the MRAG-Bench's dataset and toolkit with CCBY-4.0 license upon acceptance, to facilitate applications from both academia and industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。