arXiv:2606.03203cs.AI2026-06被引 1

首个专为医疗界面设计的自动化测试基准,验证临床软件操作可靠性。

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

论文配图:MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
图 1 · 摘自论文原文
  • 构建18个真实医疗场景,还原真实医软界面与操作流程。
  • 23个模型在真实OpenEMR上最高仅54.2%任务成功,开源模型平均仅2.5%。
  • 含临床安全维度评估,适合医疗AI、人机交互研究者使用。

计算机使用代理可自动化重复性屏幕操作,但其在医疗图形用户界面中的可靠性尚未充分验证。现有基准多聚焦通用网页或桌面任务,难以反映医疗软件特性:需领域知识、界面差异大、缺乏公开测试环境、且安全要求远超任务完成。我们提出MedCUA-Bench,一个面向临床计算机使用代理的交互式基准。涵盖10个医学领域中的18个临床场景,基于真实产品手册和开源医疗系统重建,确保界面真实性并规避版权与隐私问题。每项任务配备意图级与步骤级目标,分离临床推理与界面操作;通过确定性检查器评估任务完成度及五项临床安全维度。在23个代理中,最佳闭源模型严格成功率仅54.2%,所有模型在真实OpenEMR上均低于9%;开源模型平均仅2.5%,最高达16.2%。该基准揭示当前代理与可靠临床软件使用间的巨大差距,为未来研究提供可复现的测试平台。

原文摘要 · Abstract (English)

Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated. Existing benchmarks focus on general web or desktop tasks and underrepresent medical software, which requires domain knowledge, exhibits markedly different UI design from mainstream applications, lacks public testing environments, and demands safety validation beyond task completion. We introduce MedCUA-Bench, an interactive benchmark for clinical computer-use agents. It covers 18 clinical scenarios across 10 medical domains, reconstructed from real product manuals and open-source medical systems to capture authentic clinical interfaces while avoiding licensing and privacy constraints. Each task ships with paired intent- and step-level goals to disentangle clinical reasoning from UI execution, and is evaluated by a deterministic checker over task completion and five clinical safety dimensions. Across 23 agents, the best closed-source model reaches 54.2% strict success, while all models remain below 9% on the real OpenEMR. Open-source agents average only 2.5%, with the best reaching 16.2%. MedCUA-Bench exposes the gap between current agents and reliable clinical software use, providing a reproducible testbed for future research.

医疗AI人机交互评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。