arXiv:2501.11885cs.CL2025-01被引 9

通过循证医学流程提升大模型医疗推理能力,不需训练也能超越顶尖模型。

Med-R$^2$: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine

  • 按循证医学流程整合检索与证据推理,构建可信医疗大模型。
  • 相比基础RAG提升13.27%,优于微调策略4.55%,零训练成本。
  • 适用于临床决策支持,适合医疗AI研发与医学教育场景。

大语言模型在临床场景中展现强大能力,但现有方法存在训练成本高、数据过时等问题。依赖外部知识库虽可行,却受限于检索精度低和答案提取效果差。为此,我们提出Med-R²框架,遵循循证医学流程,高效融合证据检索、选择与推理机制,显著提升大模型在医疗任务中的问题解决能力,构建可信赖的医疗大模型。实验表明,Med-R²相较基础RAG提升13.27%,优于微调策略4.55%,且无需额外训练成本。其中,LLaMA3.1-70B + Med-R²在多项指标上超越GPT-4o(+1.05%)、Claude3.5-Sonnet(+6.14%)和DeepSeek-V3(+1.91%),有效增强大模型在医疗领域的表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have exhibited remarkable capabilities in clinical scenarios. Despite their potential, existing works face challenges when applying LLMs to medical settings. Strategies relying on training with medical datasets are highly cost-intensive and may suffer from outdated training data. Leveraging external knowledge bases is a suitable alternative, yet it faces obstacles such as limited retrieval precision and poor effectiveness in answer extraction. These issues collectively prevent LLMs from demonstrating the expected level of proficiency in mastering medical expertise. To address these challenges, we introduce Med-R^2, a novel LLM physician framework that adheres to the Evidence-Based Medicine (EBM) process, efficiently integrating retrieval mechanisms as well as the selection and reasoning processes of evidence, thereby enhancing the problem-solving capabilities of LLMs in healthcare scenarios and fostering a trustworthy LLM physician. Our comprehensive experiments indicate that Med-R^2 achieves a 13.27\% improvement over vanilla RAG methods and even a 4.55\% enhancement compared to fine-tuning strategies, without incurring additional training costs. Furthermore, we find that our LLaMA3.1-70B + Med-R$^2$ surpasses frontier models, including GPT-4o, Claude3.5-Sonnet and DeepSeek-V3 by 1.05\%, 6.14\% and 1.91\%. Med-R$^2$ effectively enhances the capabilities of LLMs in the medical domain.

医疗AI循证医学检索增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。