arXiv:2606.04127cs.CL2026-06ACL被引 1

医学问答中检索增强效果有限,模型自身能力才是关键瓶颈。

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

论文配图:When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG
图 1 · 摘自论文原文
  • 在多个模型和数据集上测试检索增强效果
  • 检索仅带来1-2分提升,远低于模型选择的影响
  • 适合关注大模型医学问答真实性能的研究者

医学问答是高风险场景,事实错误可能造成严重后果。检索增强生成(RAG)被视为有前景的解决方案,以往研究报道大型医学问答模型有显著提升。我们在此跨多种开源指令微调模型(7B至72B参数)进行大规模复现,涵盖五种模型、十组生物医学QA数据集、四种检索方法和四种检索语料库。结果发现,检索带来的提升极小且不一致,通常仅为1-2个百分点;相比之下,模型架构选择的影响远超检索器或语料库的选择。专家与普通用户来源的检索结果在多数情况下表现相近。这表明主要瓶颈并非检索质量,而是模型有效利用检索证据的能力有限。

原文摘要 · Abstract (English)

Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely viewed as a promising solution, and prior work has reported substantial gains for large medical QA models. We revisit this assumption across a broad range of open-weight instruction-tuned models spanning 7B to 72B parameters. Across five models, ten biomedical QA datasets, four retrieval methods, and four retrieval corpora, we find that retrieval yields only small and inconsistent improvements over a no-retrieval baseline, typically within 1-2 points. In contrast, the choice of backbone model has a much larger effect than the choice of retriever or corpus, and expert and layman retrieval sources perform similarly in most settings. These results suggest that the main bottleneck is not retrieval quality alone, but the model's limited ability to use retrieved evidence effectively.

医学问答RAG大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。