arXiv:2606.11702cs.CVcs.AI2026-06

构建医疗工具代理的基准,测试多步骤临床决策能力。

MedCTA: A Benchmark for Clinical Tool Agents

论文配图:MedCTA: A Benchmark for Clinical Tool Agents
图 1 · 摘自论文原文
  • 设计107个真实临床任务,包含影像、病理和报告等多模态输入。
  • 18个模型测试中,主流系统仍存在流程失败、提前终止等问题。
  • 适合评估医疗AI代理的可靠性与工具调用能力的研究者使用。

为支持临床决策,医疗AI代理需超越简单识别,具备工具检索、证据获取与整合能力。现有基准主要评估单一感知或单轮问答,难以揭示规划、工具调用与执行可靠性方面的缺陷。我们提出MedCTA,一个基于临床验证、步骤隐含的多模态临床任务基准,涵盖放射科图像、病理科切片和报告。MedCTA包含107个真实临床任务,涉及5个已部署工具,支持对工具选择、参数有效性、执行稳定性、轨迹一致性和结果质量的过程感知评估。我们对18个开源与闭源多模态模型进行评测,发现即使前沿系统在多步临床工具使用中依然脆弱:自主执行普遍出现协议错误、提前停止和错误工具调用;而黄金标准工具路径虽带来显著但不完全的性能提升。结果表明,强大的基础感知能力无法转化为临床环境中的可靠代理行为。MedCTA为审计、诊断和推进可信医疗AI代理提供了严格测试平台。数据集与评估套件可访问 https://ivul-kaust.github.io/MedCTA/

原文摘要 · Abstract (English)

To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Existing benchmarks largely evaluate isolated perception or single-turn question answering, and therefore provide limited visibility into failures of planning, tool recruitment, and rollout reliability. We introduce MedCTA, a benchmark for evaluating medical tool agents on clinician-validated, step-implicit tasks grounded in realistic multimodal clinical inputs, including radiology images, pathology slides, and reports. MedCTA comprises 107 real-world clinical tasks with clinician-verified executable trajectories over 5 deployed tools, and supports process-aware evaluation of tool selection, argument validity, execution stability, trajectory fidelity, and outcome quality. We benchmark 18 open- and closed-source multimodal models and find that even frontier systems remain brittle in multi-step clinical tool use: autonomous rollouts are dominated by protocol failures, premature stopping, and incorrect tool recruitment, while gold-standard tool routing yields large but still incomplete gains. These results show that strong backbone perception does not translate into reliable agentic behavior in clinical settings. MedCTA provides a rigorous testbed for auditing, diagnosing, and advancing trustworthy medical AI agents. The dataset and evaluation suite are available at https://ivul-kaust.github.io/MedCTA/

医疗AI工具代理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。