测试AI代理在复杂医疗流程中的端到端自动化能力
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

- 构建多角色、长周期、政策密集的医疗工作流基准
- 最佳代理仅完成28.0%任务,严格通过率不足20%
- 适合研究企业级复杂决策系统与AI协作的学者
当前基准普遍缺乏对真实医疗操作自动化所需的三大能力的评估:政策密度(需遵循超1290份医疗、保险与运营规则)、多角色协同(单任务需跨角色交接)以及多方交互(如多轮同行评审与患者沟通)。我们提出χ-Bench,涵盖医疗机构预授权、支付方利用管理与护理管理三个领域,每个任务在高保真模拟环境中,通过87个MCP工具操控20个医疗应用,以完成临床案例的闭环处理。代理需依据超过1290份管理护理操作手册进行决策,通过工具调用与撰写角色文档推进流程。在30种代理配置中,最高完成率为28.0%,严格通过率(pass^3)低于20%;单会话执行全部任务时性能骤降至3.8%。这些结果表明,在其他政策密集、角色复合且不可逆的企业场景中,类似能力鸿沟可能普遍存在。
原文摘要 · Abstract (English)
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce $χ$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill. Across 30 agent harness/models configurations, the best agent resolves only 28.0% of tasks, no agent clears 20% on strict pass^3, and executing all tasks in a single session slumps the performance to 3.8%. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。