arXiv:2505.05283cs.SEcs.AI2025-05综述被引 26

系统梳理代码大模型在软件开发全周期的评测现状与不足

Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents

  • 按软件开发生命周期分层分析178个评测基准
  • 六成评测集中于编码阶段,需求与设计阶段严重缺失
  • 多数缺乏防数据泄露机制,可能夸大模型真实能力

代码大语言模型(CodeLLMs)和智能体正被广泛应用于覆盖整个软件开发生命周期(SDLC)的复杂工程任务中。评测是严格评估其能力的关键。然而,尽管重要性日益凸显,目前仍缺乏从SDLC视角对这些评测基准的全面综述。为填补这一空白,我们提出一种分层分析框架,系统回顾了461篇论文中的178个评测基准,全面刻画其在SDLC各阶段的分布特征。研究发现,当前评测存在显著失衡:约61%聚焦于软件实现阶段,而需求工程和软件设计阶段分别仅有5%和3%。此外,多数评测缺乏有效的反污染策略,导致数据泄露风险高,可能引发性能评估虚高。最后,我们识别出当前研究中的关键开放挑战,并提出未来方向,以缩小代码大模型理论能力与实际应用效果之间的差距。

原文摘要 · Abstract (English)

Code large language models (CodeLLMs) and agents are increasingly being integrated into complex software engineering tasks spanning the entire Software Development Life Cycle (SDLC). Benchmarking is critical for rigorously evaluating these capabilities. However, despite their growing significance, there remains a lack of comprehensive reviews that examine these benchmarks from an SDLC perspective. To bridge this gap, we propose a tiered analysis framework to systematically review 178 benchmarks from 461 papers, comprehensively characterizing them from the perspective of the SDLC. Our findings reveal a notable imbalance in the coverage of current benchmarks, with approximately 61\% focused on the software implementation phase in SDLC, while requirements engineering and software design phases receive minimal attention at only 5\% and 3\%, respectively. % Additionally, anti-contamination strategies are largely absent from current benchmarks, leading to an increased risk of data leakage. Furthermore, current benchmarks lack effective anti-contamination strategies, posing significant risks of data leakage and potentially inflated performance assessments. Finally, we identify key open challenges in current research and outline future directions to narrow the gap between the theoretical capabilities of CodeLLMs and agents and their practical effectiveness in real-world scenarios.

代码生成评测基准软件工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。