arXiv:2512.10996cs.CLcs.AI2025-12被引 4

用语义搜索和微调提升医学生物问答准确率

MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA

  • 融合语义与关键词检索,精准定位生物医学文档
  • 在多个基准数据集上超越GPT-4o等先进模型
  • 适合需要高精度医学问答的研究者使用

近年来,检索增强生成(RAG)显著提升了大语言模型(LLMs)在复杂问答任务中的表现。本文提出MedBioRAG,一种结合语义与词汇搜索、文档检索及监督微调的生物医学问答模型。该模型高效检索并排序相关生物医学文献,实现精准且上下文相关的回答生成。我们在NFCorpus、TREC-COVID、MedQA、PubMedQA和BioASQ等多个基准数据集上评估了MedBioRAG,涵盖文本检索、封闭式问答和长文本问答任务。实验结果表明,其在所有任务中均优于此前最先进(SoTA)模型及GPT-4o基础模型。尤其在文档检索中提升了NDCG和MRR指标,在封闭式问答中准确率更高,在长文本问答中获得更优的ROUGE分数。研究验证了基于语义搜索的检索与LLM微调在生物医学应用中的有效性。

原文摘要 · Abstract (English)

Recent advancements in retrieval-augmented generation (RAG) have significantly enhanced the ability of large language models (LLMs) to perform complex question-answering (QA) tasks. In this paper, we introduce MedBioRAG, a retrieval-augmented model designed to improve biomedical QA performance through a combination of semantic and lexical search, document retrieval, and supervised fine-tuning. MedBioRAG efficiently retrieves and ranks relevant biomedical documents, enabling precise and context-aware response generation. We evaluate MedBioRAG across text retrieval, close-ended QA, and long-form QA tasks using benchmark datasets such as NFCorpus, TREC-COVID, MedQA, PubMedQA, and BioASQ. Experimental results demonstrate that MedBioRAG outperforms previous state-of-the-art (SoTA) models and the GPT-4o base model in all evaluated tasks. Notably, our approach improves NDCG and MRR scores for document retrieval, while achieving higher accuracy in close-ended QA and ROUGE scores in long-form QA. Our findings highlight the effectiveness of semantic search-based retrieval and LLM fine-tuning in biomedical applications.

医学问答检索增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。