arXiv:2501.12746cs.CLcs.AI2025-01被引 5

让小模型学会分析医学证据,回答更准更可靠。

EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering

  • 用6600万参数小模型显式学习证据支持度、逻辑关联和摘要。
  • 在参考质量与准确率上超越80亿参数大模型的RAG方法。
  • 适合资源受限场景下的精准医学问答任务。

在应对生物医学领域的专业问题时,人类通常需收集多条证据并进行多维度分析以给出高质量答案。当前基于大语言模型的问答方法缺乏对证据分析的明确定义与学习过程,易导致错误传播与幻觉。尽管增大模型参数可缓解问题,但会带来训练与部署资源压力。本研究提出EvidenceMap,使一个极小的预训练语言模型(约30亿参数)能显式学习生物医学证据的多重特性,包括支持性评估、逻辑关联与内容摘要,从而隐式引导小生成模型生成文本回答。实验表明,通过仅6600万参数模型微调学习证据分析,其在参考依据质量与准确率上分别比使用80亿参数大模型的RAG方法高出19.9%与5.7%。

原文摘要 · Abstract (English)

When addressing professional questions in the biomedical domain, humans typically acquire multiple pieces of information as evidence and engage in multifaceted analysis to provide high-quality answers. Current LLM-based question answering methods lack a detailed definition and learning process for evidence analysis, leading to the risk of error propagation and hallucinations while using evidence. Although increasing the parameter size of LLMs can alleviate these issues, it also presents challenges in training and deployment with limited resources. In this study, we propose EvidenceMap, which aims to enable a tiny pre-trained language model to explicitly learn multiple aspects of biomedical evidence, including supportive evaluation, logical correlation and content summarization, thereby latently guiding a small generative model (around 3B parameters) to provide textual responses. Experimental results demonstrate that our method, learning evidence analysis by fine-tuning a model with only 66M parameters, exceeds the RAG method with an 8B LLM by 19.9% and 5.7% in reference-based quality and accuracy, respectively.

医学问答小模型证据分析RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。