测试AI运维代理在真实系统中的表现,发现多数无法稳定完成任务。
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

- 构建跨系统层级与生命周期的基础设施评估基准
- 15种配置中平均得分40%~88%,顶尖配置仍常失败
- 揭示代理常留隐患:短期达标但留下不持久变更和状态污染
现代计算基础设施管理因复杂性持续上升而愈发困难。人工智能代理的进展为自动化运维带来契机,但其应对真实世界复杂性的能力尚不明确。本文提出InfraBench,一个涵盖完整系统栈与运维生命周期的基准套件,支持细粒度风险评估。对15种代理模型配置的实验表明,即使最强配置也无法在所有任务中获得满分,平均有效得分介于40%至88%之间(各配置标准误差6-12个百分点)。重复三次任务测试显示,顶级配置仍仅能成功部分尝试;逐检查项评分揭示普遍失效模式:代理常达成短期目标,却遗留非持久性变更、分布式不变量破坏、安全隐患及未清理状态。InfraBench(含实时排行榜、任务与评估工具)已公开,网址为 infraben.ch。
原文摘要 · Abstract (English)
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。