arXiv:2508.14880cs.CL2025-08被引 22

用医学知识图谱和私有检索引擎,让小模型在医疗研究上超越大公司系统。

MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework

  • 基于医学知识图谱生成多跳问答数据,构建复杂推理路径。
  • 整合专用医疗检索工具,平均每个任务触发4.2次工具调用。
  • 开源小模型在12个医学领域表现超大型闭源系统,适合医疗科研人员。

大型语言模型驱动的智能体在多领域展现强大能力,尤其在复杂信息检索与综合任务中表现优异。然而,通用深度研究智能体在医疗领域面临显著挑战,现有领先闭源系统在复杂医学基准测试中准确率有限。主要瓶颈在于:(1) 模型缺乏足够的密集医学知识以支持临床推理;(2) 框架受限于缺乏针对医疗场景的专用检索工具。本文提出MedResearcher-R1,通过两项核心创新应对上述问题。首先,设计一种基于医学知识图谱的数据合成框架,从罕见医学实体周围的子图中提取最长链,生成超过2100个多跳问答轨迹,覆盖12个医学专科。其次,集成自研私有医疗检索引擎与通用工具,实现精准医疗信息融合。采用两阶段训练范式(监督微调+在线强化学习,复合奖励机制),所提出的MedResearcher-R1-32B模型在医学基准上取得新SOTA,同时保持在通用深度研究任务上的竞争力。结果表明,通过架构、工具与训练数据的领域专项优化,小型开源模型可在专业领域超越更大规模的闭源系统。

原文摘要 · Abstract (English)

Recent developments in Large Language Model (LLM)-based agents have shown impressive capabilities spanning multiple domains, exemplified by deep research systems that demonstrate superior performance on complex information-seeking and synthesis tasks. While general-purpose deep research agents have shown impressive capabilities, they struggle significantly with medical domain challenges, as evidenced by leading proprietary systems achieving limited accuracy on complex medical benchmarks. The key limitations are: (1) the model lacks sufficient dense medical knowledge for clinical reasoning, and (2) the framework is constrained by the absence of specialized retrieval tools tailored for medical contexts. We present a medical deep research agent that addresses these challenges through two core innovations. First, we develop a novel data synthesis framework using medical knowledge graphs, extracting the longest chains from subgraphs around rare medical entities to generate complex multi-hop question-answer pairs. Second, we integrate a custom-built private medical retrieval engine alongside general-purpose tools, enabling accurate medical information synthesis. Our approach generates 2100+ diverse trajectories across 12 medical specialties, each averaging 4.2 tool interactions. Through a two-stage training paradigm combining supervised fine-tuning and online reinforcement learning with composite rewards, our MedResearcher-R1-32B model demonstrates exceptional performance, establishing new state-of-the-art results on medical benchmarks while maintaining competitive performance on general deep research tasks. Our work demonstrates that strategic domain-specific innovations in architecture, tool design, and training data construction can enable smaller open-source models to outperform much larger proprietary systems in specialized domains.

医疗AI大模型知识图谱研究智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。