打造能统管多种任务的AI程序员,提升自动化开发能力。
Unified Software Engineering Agent as AI Software Engineer
- 构建统一代理架构,整合编码、测试、修复等多类能力。
- 在1271个项目级任务中表现优于现有通用代理。
- 适合未来人机协同开发团队,推动AI工程师落地。
大型语言模型(LLM)技术的发展带来了自动化编程的期待。然而,软件工程远不止编码,还涉及维护与演化等多类活动。在此背景下,基于LLM的智能体因其可自主调用外部工具而受到关注。但这类智能体是否等同于AI软件工程师?本文通过构建统一软件工程智能体(USEagent)来回答此问题。不同于针对特定任务(如测试、调试、修复)的专用智能体,本研究目标是打造一个能协调并处理多种能力的统一智能体,具备应对复杂开发场景的能力,如修复不完整补丁、新增功能或接手他人代码。我们设想USEagent是未来AI软件工程师的雏形,可作为人机协作开发团队的一员。为评估其效能,我们构建了统一软件工程基准(USEbench),融合SWE-bench、SWT-bench和REPOCOD等多个基准的任务。在包含1,271个项目级软件工程任务的评估中,USEagent相较OpenHands CodeActAgent等现有通用代理展现出更优性能。然而,在部分编码任务上仍存在能力差距,为未来AI软件工程师的发展提供了方向。
原文摘要 · Abstract (English)
The growth of Large Language Model (LLM) technology has raised expectations for automated coding. However, software engineering is more than coding and is concerned with activities including maintenance and evolution of a project. In this context, the concept of LLM agents has gained traction, which utilize LLMs as reasoning engines to invoke external tools autonomously. But is an LLM agent the same as an AI software engineer? In this paper, we seek to understand this question by developing a Unified Software Engineering agent or USEagent. Unlike existing work which builds specialized agents for specific software tasks such as testing, debugging, and repair, our goal is to build a unified agent which can orchestrate and handle multiple capabilities. This gives the agent the promise of handling complex scenarios in software development such as fixing an incomplete patch, adding new features, or taking over code written by others. We envision USEagent as the first draft of a future AI Software Engineer which can be a team member in future software development teams involving both AI and humans. To evaluate the efficacy of USEagent, we build a Unified Software Engineering bench (USEbench) comprising of myriad tasks such as coding, testing, and patching. USEbench is a judicious mixture of tasks from existing benchmarks such as SWE-bench, SWT-bench, and REPOCOD. In an evaluation on USEbench consisting of 1,271 repository-level software engineering tasks, USEagent shows improved efficacy compared to existing general agents such as OpenHands CodeActAgent. There exist gaps in the capabilities of USEagent for certain coding tasks, which provides hints on further developing the AI Software Engineer of the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。