arXiv:2604.09937cs.AI2026-04被引 8

首个评估医疗行政代理的基准,揭示现有模型在真实流程中可靠性不足。

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks

  • 构建四类真实医疗界面环境与135项专家任务,拆解为1698个可验证子任务。
  • 最佳代理仅36.3%任务成功率,子任务最高达82.8%,暴露端到端能力短板。
  • 适合关注医疗自动化、智能代理评估的研究者与从业者参考。

医疗行政支出每年超过1万亿美元,是大语言模型驱动计算机使用代理(CUAs)的重要目标。尽管临床应用受到广泛关注,但缺乏针对端到端行政工作流的评估基准。为此,我们提出HealthAdminBench,包含四种真实图形界面环境:电子病历系统(EHR)、两个保险商门户和一个传真系统,涵盖135项由专家定义的任务,覆盖三大行政任务类型:预先授权、申诉与拒赔管理、耐用医疗设备(DME)订单处理。每项任务被分解为细粒度、可验证的子任务,共形成1,698个评估点。我们在多种提示与观察设置下评估七种代理配置,发现尽管子任务表现良好,但端到端可靠性仍低:表现最佳的代理(Claude Opus 4.6 CUA)任务成功率为36.3%,GPT-5.4 CUA达到最高子任务成功率(82.8%)。结果揭示了当前代理能力与真实行政流程需求之间的显著差距。HealthAdminBench为评估医疗行政流程自动化进展提供了严格基准。

原文摘要 · Abstract (English)

Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While clinical applications of LLMs have received significant attention, no benchmark exists for evaluating CUAs on end-to-end administrative workflows. To address this gap, we introduce HealthAdminBench, a benchmark comprising four realistic GUI environments: an EHR, two payer portals, and a fax system, and 135 expert-defined tasks spanning three administrative task types: Prior Authorization, Appeals and Denials Management, and Durable Medical Equipment (DME) Order Processing. Each task is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluation points. We evaluate seven agent configurations under multiple prompting and observation settings and find that, despite strong subtask performance, end-to-end reliability remains low: the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent). These results reveal a substantial gap between current agent capabilities and the demands of real-world administrative workflows. HealthAdminBench provides a rigorous foundation for evaluating progress toward safe and reliable automation of healthcare administrative workflows.

医疗自动化代理评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。