首个评估医学指南深度证据整合能力的基准,推动大模型向专家级医疗决策迈进。
MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

- 以真实临床指南为专家参考,构建多步骤证据整合评测任务
- 通过5130+原子论点验证证据精度,1200+条任务自适应评分标准全面评估质量
- 揭示当前大模型在指南生成与证据整合上距专家水平仍有显著差距
近期深度研究系统使大语言模型能够检索、综合并推理大规模外部知识。在医学领域,制定临床指南高度依赖此类深度证据整合能力。然而,现有基准无法在需多步证据整合与专家级判断的真实工作流中有效评估该能力。为此,我们提出MedProbeBench,首个利用高质量临床指南作为专家级参考的基准。临床指南以其严格的中立性与可验证性标准代表医学专业性的巅峰,对深度研究智能体构成重大挑战。评估方面,我们设计MedProbe-Eval框架:(1) 全面评分体系,包含1200+任务自适应评分标准,实现质量全面评估;(2) 细粒度证据验证,基于5130+原子论点严格检验证据精确性。对17个LLM和深度研究代理的评估揭示了其在证据整合与指南生成方面的关键短板,凸显当前能力与专家级临床指南制定之间仍存在显著差距。
原文摘要 · Abstract (English)
Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medicine, developing clinical guidelines critically depends on such deep evidence integration. However, existing benchmarks fail to evaluate this capability in realistic workflows requiring multi-step evidence integration and expert-level judgment. To address this gap, we introduce MedProbeBench, the first benchmark leveraging high-quality clinical guidelines as expert-level references. Medical guidelines, with their rigorous standards in neutrality and verifiability, represent the pinnacle of medical expertise and pose substantial challenges for deep research agents. For evaluation, we propose MedProbe-Eval, a comprehensive evaluation framework featuring: (1) Holistic Rubrics with 1,200+ task-adaptive rubric criteria for comprehensive quality assessment, and (2) Fine-grained Evidence Verification for rigorous validation of evidence precision, grounded in 5,130+ atomic claims. Evaluation of 17 LLMs and deep research agents reveals critical gaps in evidence integration and guideline generation, underscoring the substantial distance between current capabilities and expert-level clinical guideline development. Project: https://github.com/uni-medical/MedProbeBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。