arXiv:2512.14896cs.CLcs.AI2025-12被引 1

用外部药学知识提升大模型在执业药师题库上的答题准确率。

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

  • 构建三步检索增强生成流程,外挂结构化药学证据优化提示词。
  • 所有模型准确率提升7至21个百分点,小模型提升更显著。
  • 无需修改模型架构,适合需高精度药学问答的AI应用。

本研究评估了大语言模型(LLM)在执业药师风格问答任务中的表现,并开发了一种外部知识融合方法以提升准确性。我们使用包含141道题的药学数据集,对参数量从80亿到700亿以上的十种LLM进行了基准测试,未经过任何修改的基线准确率介于46%至92%之间,其中GPT-5(92%)和o3(89%)表现最佳,而较小的开源模型表现明显较低。随后,我们提出DrugRAG,一种三步检索增强生成(RAG)流程,通过检索结构化、基于证据的药物信息,并将其作为上下文注入模型提示,全程外部运行,无需更改模型架构或参数。DrugRAG在全部五种评估模型中均提升了准确率,增幅为7至21个百分点(如:Gemma 3 27B从61.0%升至71%,Llama 3.1 8B从46%升至67%)。McNemar检验显示,小模型和中等规模开源模型的改进具有统计显著性。结果表明,通过DrugRAG集成结构化外部药物知识,可在不修改底层模型的前提下提升其在药学问答任务中的表现,为构建基于证据的药学类AI应用提供了实用管道。

原文摘要 · Abstract (English)

In our study, we evaluated large language model (LLM) performance on pharmacy licensure-style question-answering tasks and developed an external knowledge integration method to improve accuracy. We benchmarked ten LLMs with varying parameter sizes (8 billion to 70+ billion) using a 141-question pharmacy dataset, measuring baseline accuracy without modification. Baseline performance ranged from 46% to 92%, with GPT-5 (92%) and o3 (89%) achieving the highest scores, while smaller open-source models showed substantially lower performance. We then developed DrugRAG, a three-step retrieval-augmented generation (RAG) pipeline that retrieves structured, evidence-based drug information and augments model prompts with contextual pharmacological evidence, operating externally and requiring no changes to model architecture or parameters. DrugRAG increased accuracy across all five evaluated models, with gains ranging from 7 to 21 percentage points (e.g., Gemma 3 27B: 61.0% to 71%, Llama 3.1 8B: 46% to 67%). McNemar analyses demonstrated statistically significant paired improvements primarily in smaller and mid-sized open-source models. These findings demonstrate that integrating structured external drug knowledge via DrugRAG can improve LLM performance on pharmacy-focused question-answering tasks without modifying the underlying models, providing a practical pipeline for enhancing evidence-based pharmacy-focused AI applications.

药学AIRAG大模型增强知识注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。