打造医疗深度搜索代理,提升中文医学领域智能检索能力
QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence

- 融合医学知识图谱与实时网络探索构建长时序医疗搜索数据
- 采用两阶段微调与强化学习,显著增强规划与工具调用能力
- 联合专家构建权威评测基准,适合医疗AI研究者使用
随着代理型基础模型的发展,如何进一步提升其在垂直领域的性能成为关键挑战。为此,基于通义DeepResearch这一强大代理型基础模型,我们聚焦中文医疗深度搜索场景,提出QuarkMedSearch,系统性地探索涵盖医疗多跳数据构建、训练策略与评估基准的全流程方法,以进一步突破垂直领域性能上限。针对医疗领域深度搜索训练数据稀缺问题,结合大规模医学知识图谱与实时在线探索,构建长时序医疗深度搜索训练数据;在后训练阶段,采用两阶段SFT与强化学习策略,逐步提升模型在规划、工具调用与反思能力方面的表现,同时保持搜索效率;在评估方面,联合医学专家通过严格人工验证构建QuarkMedSearch基准。实验结果表明,QuarkMedSearch在同规模开源模型中于QuarkMedSearch基准上达到最先进水平,同时在通用基准上仍具较强竞争力。
原文摘要 · Abstract (English)
As agentic foundation models continue to evolve, how to further improve their performance in vertical domains has become an important challenge. To this end, building upon Tongyi DeepResearch, a powerful agentic foundation model, we focus on the Chinese medical deep search scenario and propose QuarkMedSearch, systematically exploring a full-pipeline approach spanning medical multi-hop data construction, training strategies, and evaluation benchmarks to further push and assess its performance upper bound in vertical domains. Specifically, for data synthesis, to address the scarcity of deep search training data in the medical domain, we combine a large-scale medical knowledge graph with real-time online exploration to construct long-horizon medical deep search training data; for post-training, we adopt a two-stage SFT and RL training strategy that progressively enhances the model's planning, tool invocation, and reflection capabilities required for deep search, while maintaining search efficiency; for evaluation, we collaborate with medical experts to construct the QuarkMedSearch Benchmark through rigorous manual verification. Experimental results demonstrate that QuarkMedSearch achieves state-of-the-art performance among open-source models of comparable scale on the QuarkMedSearch Benchmark, while also maintaining strong competitiveness on general benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。