打造高保真故障场景,评估AI运维代理真实表现
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

- 构建可扩展的实时系统环境,模拟多层故障与噪声
- 涵盖90个真实复杂SRE问题,代理表现差异达40%
- 适合研究者和工程师测试AI运维工具在生产级场景的能力
AI代理在诊断和缓解生产系统故障中的应用日益广泛,即智能运维(SRE)。现有SRE基准测试局限于过于简化的任务,且因定制化设计难以扩展。本文提出SREGym,一个高保真度的SRE代理评估基准。SREGym基于真实云原生系统架构搭建实时运行环境,通过故障注入器模拟高保真故障场景。其通过模拟(1)多层级的广泛故障,(2)各类环境噪声,(3)如亚稳态故障和相关性故障等多样化故障模式,刻画生产环境的复杂性。SREGym采用模块化、可扩展的框架,跨系统栈协同编排故障与噪声注入器。当前包含90个真实且具挑战性的SRE问题。我们使用SREGym评估前沿代理,发现其在应对不同故障类型时表现差异显著,端到端结果最高相差达40%。SREGym为开源项目,持续维护,已被研究人员与从业者广泛使用。
原文摘要 · Abstract (English)
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to oversimplistic SRE tasks and are unfortunately hard to extend due to bespoke designs. We present SREGym, a high-fidelity benchmark for SRE agents. SREGym exposes a live system environment built atop real-world cloud-native system stacks, where high-fidelity failure scenarios are simulated through fault injectors. SREGym models the complexity of production environments by simulating (1) a wide range of faults at different layers, (2) various ambient noises, and (3) diverse failure modes such as metastable failures and correlated failures. SREGym is architected as a modular, extensible framework that orchestrates fault and noise injectors across stacks. SREGym currently includes 90 realistic, challenging SRE problems. We use SREGym to evaluate frontier agents and show that their capabilities varies significantly in addressing different kinds of failures, with up to 40% differences in end-to-end results. SREGym is actively maintained as an open-source project and has been used by researchers and practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。