首个评估AI在完整软件开发运维流程中表现的基准测试
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
- 构建涵盖构建配置、监控、问题修复和测试生成的全流程评测环境
- 包含30多个项目中的700多个真实任务,覆盖Java和Go语言
- 揭示当前AI代理在复杂运维任务上存在显著短板,适合研究自动化运维者参考
尽管人工智能代理在代码生成和软件问题修复方面展现出卓越能力,但其在完整软件开发运维(DevOps)周期中的表现仍不明确。与纯代码生成不同,真实世界的DevOps流程涉及大规模项目分析、动态程序行为理解、领域专用工具使用及序列决策。现有基准测试多聚焦孤立问题,缺乏支持完整流程的环境与工具接口。我们提出DevOps-Gym,首个面向核心DevOps工作流的端到端基准:构建与配置、监控、问题修复和测试生成。该基准包含从30余个真实项目中收集的700多个任务,覆盖Java和Go语言。我们设计了半自动数据收集机制,结合严谨且非平凡的专家投入以确保任务覆盖率与质量。对前沿模型与代理的评估发现:它们在Java和Go的问题修复与测试生成上表现不佳,且无法处理监控、构建与配置等新任务。这些结果凸显了在全链路自动化运维中开展关键研究的迫切性。
原文摘要 · Abstract (English)
Even though demonstrating extraordinary capabilities in code generation and software issue resolving, AI agents' capabilities in the full software DevOps cycle are still unknown. Different from pure code generation, handling the DevOps cycle in real-world software, including developing, deploying, and managing, requires analyzing large-scale projects, understanding dynamic program behaviors, leveraging domain-specific tools, and making sequential decisions. However, existing benchmarks focus on isolated problems and lack environments and tool interfaces for DevOps. We introduce DevOps-Gym, the first end-to-end benchmark for evaluating AI agents across core DevOps workflows: build and configuration, monitoring, issue resolving, and test generation. DevOps-Gym includes 700+ real-world tasks collected from 30+ projects in Java and Go. We develop a semi-automated data collection mechanism with rigorous and non-trivial expert efforts in ensuring the task coverage and quality. Our evaluation of state-of-the-art models and agents reveals fundamental limitations: they struggle with issue resolving and test generation in Java and Go, and remain unable to handle new tasks such as monitoring and build and configuration. These results highlight the need for essential research in automating the full DevOps cycle with AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。