arXiv:2505.14963cs.CL2025-05被引 26

首个测试医疗大模型多跳检索与推理的基准,揭示其临床应用能力严重不足。

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use

  • 构建真实临床场景的1000+人工标注问题,测试多源知识融合能力
  • 前沿智能体在复杂任务中准确率最低仅10%,远低于临床需求
  • 适合评估医疗AI工具的真实可靠性,推动模型改进

大型语言模型(LLMs)被寄望用于临床决策支持,但安全的临床推理需在严格精度约束下整合试验、原始研究、监管文件和成本数据等异构知识库。现有评估常依赖合成提示,将任务简化为单跳事实查询,或混淆推理与开放式生成,难以反映实际效用。为此,我们提出MedBrowseComp,首个系统性测试智能体从实时领域知识库中可靠检索并综合多跳医学事实的基准。该基准包含1000多个仿照临床情境的人工标注问题,要求从业者在碎片化或冲突信息中得出最新结论。对前沿代理系统的测试显示,性能短板低至10%,暴露出当前LLM能力与临床严谨性要求之间的巨大差距。MedBrowseComp为可靠的医疗信息检索提供了清晰的测试平台,并为未来模型与工具链升级设定了具体目标。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bases -- trials, primary studies, regulatory documents, and cost data -- under strict accuracy constraints. Existing evaluations often rely on synthetic prompts, reduce the task to single-hop factoid queries, or conflate reasoning with open-ended generation, leaving their real-world utility unclear. To close this gap, we present MedBrowseComp, the first benchmark that systematically tests an agent's ability to reliably retrieve and synthesize multi-hop medical facts from live, domain-specific knowledge bases. MedBrowseComp contains more than 1,000 human-curated questions that mirror clinical scenarios where practitioners must reconcile fragmented or conflicting information to reach an up-to-date conclusion. Applying MedBrowseComp to frontier agentic systems reveals performance shortfalls as low as ten percent, exposing a critical gap between current LLM capabilities and the rigor demanded in clinical settings. MedBrowseComp therefore offers a clear testbed for reliable medical information seeking and sets concrete goals for future model and toolchain upgrades. You can visit our project page at: https://moreirap12.github.io/mbc-browse-app/

医疗AI多跳推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。