将代码库自动转化为可协作的智能软件代理,推动网络自治化发展
SW-$A^2$-Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web

- 用编码代理自动把代码库转为可运行的软件代理
- 生成的代理能准确还原源代码功能且支持多代理协同
- 首个针对软件代理生成的基准测试,适合研究者与开发者参考
自主软件代理正成为实现网络自治的新范式,但其规模受限于可用代理数量不足。为此,本文研究通过编码代理将现有代码仓库自动转化为自主软件代理的任务,分解生成流程并识别关键技术瓶颈。提出首个专门针对软件代理生成的基准测试——SW-$A^2$-Bench,不仅能评估是否成功生成代理,还检验其对源代码的忠实度以及在多代理工作流中的互操作性。实验表明,该方法有效激活代码仓库的功能,并支持多代理间的协同合作。本工作为软件代理生成提供了标准化评估框架,有望推动阿吉恩网络(Agentic Web)的规模化发展。
原文摘要 · Abstract (English)
The Agentic Web is emerging as a paradigm in which autonomous software agents interact with online resources and with each other to accomplish user goals. However, the capacity of Agentic Web is still limited by insufficient autonomous software agent population, which has become a crucial challenge for scaling Agentic Web. In order to alleviate this, we study the task of automatically converting existing code repositories into autonomous software agents via coding agents, decompose the process into critical stages, and identify key technical hurdles. To systematically evaluate this capability, we propose SoftWare Agent generation for Agentic Web Bench (SW-$A^2$-Bench), the first benchmark designed for software agent generation. SW-$A^2$-Bench evaluates not only whether software agents can be generated, but also whether generated software agents are faithful to the source repositories and interoperable with other agents in multi-agent workflows. Our experiments demonstrate that our approach effectively activates the functional capabilities of code repositories and enables interoperable multi-agent collaboration in Agentic Web. We believe that this work will provide a standardized evaluation for software agent generation and will contribute to the future of scaling the capacity of Agentic Web.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。