arXiv:2510.09721cs.SEcs.CL2025-10综述被引 21

首份系统综述LLM智能体在软件工程中的评测与解决方案

A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System

  • 按提示、微调、智能体三类梳理解决方案
  • 关联50+评测基准,覆盖代码生成、修复等任务
  • 揭示从简单提示到多智能体协作的演进路径

大型语言模型(LLMs)在软件工程中的融合推动了从传统规则系统向自主智能体系统的转变,能够解决复杂问题。然而,系统性进展受限于对评测基准与解决方案之间关系的不完整理解。本文首次对LLM驱动的软件工程进行整体分析,涵盖150余篇近期论文,提出双维度分类体系:(1)解决方案分为提示式、微调式和智能体式;(2)评测基准包括代码生成、翻译、修复等任务。分析显示,发展轨迹已从简单提示工程演变为包含规划、推理、记忆机制和工具增强的复杂智能体系统。为此,我们构建统一工作流框架,展示从任务定义到交付成果的全过程,并阐明不同方案应对不同复杂度的能力。不同于以往聚焦单一方向的综述,本工作将50多个基准与对应策略连接,助力研究者根据评估需求选择最优方法。同时识别出关键研究空白,如多智能体协作、自演化系统与形式化验证集成,并提出未来方向。相关论文持续更新于GitHub仓库:https://github.com/lisaGuojl/LLM-Agent-SE-Survey。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into software engineering has driven a transition from traditional rule-based systems to autonomous agentic systems capable of solving complex problems. However, systematic progress is hindered by a lack of comprehensive understanding of how benchmarks and solutions interconnect. This survey addresses this gap by providing the first holistic analysis of LLM-powered software engineering, offering insights into evaluation methodologies and solution paradigms. We review over 150 recent papers and propose a taxonomy along two key dimensions: (1) Solutions, categorized into prompt-based, fine-tuning-based, and agent-based paradigms, and (2) Benchmarks, including tasks such as code generation, translation, and repair. Our analysis highlights the evolution from simple prompt engineering to sophisticated agentic systems incorporating capabilities like planning, reasoning, memory mechanisms, and tool augmentation. To contextualize this progress, we present a unified pipeline illustrating the workflow from task specification to deliverables, detailing how different solution paradigms address various complexity levels. Unlike prior surveys that focus narrowly on specific aspects, this work connects 50+ benchmarks to their corresponding solution strategies, enabling researchers to identify optimal approaches for diverse evaluation criteria. We also identify critical research gaps and propose future directions, including multi-agent collaboration, self-evolving systems, and formal verification integration. This survey serves as a foundational guide for advancing LLM-driven software engineering. We maintain a GitHub repository that continuously updates the reviewed and related papers at https://github.com/lisaGuojl/LLM-Agent-SE-Survey.

智能体系统软件工程评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。