构建真实医疗场景的智能体评估基准,测试复杂医疗任务完成能力。
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

- 设计54个跨7类的真实医疗流程任务,模拟完整临床工作流。
- 顶尖模型成功率仅42%,显示当前AI在医疗决策中仍存在显著短板。
- 适合评估医疗AI代理的推理与多步规划能力,尤其关注影像与复杂搜索任务。
随着AI智能体在复杂、长周期推理方面日益强大,对其进行全面评估对于衡量其向真实医疗应用演进的进展至关重要。我们提出了HealthAgentBench,一个包含54个智能体医疗任务的基准套件,覆盖7个类别,每类均有独特环境。该套件涵盖患者诊疗全流程中的多样化工作流和多模态数据。每个任务均模拟端到端临床流程:给定极少指令,智能体需探索原始医疗数据,在复杂环境中操作,并执行超越简单提示的多步骤解决方案。最终以任务成功率作为整体性能的单一可解释指标。在HealthAgentBench上评估前沿智能体,发现总体任务成功率仍然较低,凸显该基准的挑战性。最强且成本最低的模型Codex GPT-5.5仅达到约42%的成功率。除总体表现外,该基准揭示了不同任务类别中的优劣差异:前沿智能体在自动构建电子病历数据研究管道方面展现潜力,但医学影像任务尤其困难,尤其是Claude Code模型;而Codex GPT-5.5则初现能力。结合大规模搜索空间与组合推理需求的任务对所有现有智能体而言仍属难题。这些结果表明,HealthAgentBench提供了一个具有挑战性且现实的评估基准,未来仍有巨大提升空间。我们已将该基准开源发布于https://github.com/microsoft/HealthAgentBench。
原文摘要 · Abstract (English)
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。