arXiv:2605.25536cs.SEcs.AI2026-05综述

综述大模型代码生成研究趋势与挑战,揭示真实场景应用的短板。

A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions

论文配图:A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions
图 1 · 摘自论文原文
  • 整合30项二次研究,系统分析大模型代码生成现状
  • 基准测试准确率高,但真实场景泛化能力弱,效率与鲁棒性差
  • 适合关注代码生成评估、模型改进与工程落地的研究者

大型语言模型(LLMs)在软件工程中的代码生成任务(CGTs)中日益普及。尽管已有研究结果乐观,但其广泛应用的综合影响及在真实开发中的集成情况仍不明确,现有综述类研究在此方面覆盖不足。本综述整合了关于基于大模型的代码生成任务的二次证据,涵盖发表格局、实际效果、应用场景、集成挑战与未来研究方向。采用系统综述方法,在相关数字图书馆检索,并辅以前后滚动检索和筛选流程。研究质量通过评估,数据提取可靠性通过评分者一致性检验。证据融合使用SWEBOK知识领域与HELM框架进行综合分析。共识别出2017至2025年间发表的30项二次研究,自2023年起呈快速增长趋势。基准测试显示准确性较高,但在真实世界中的泛化能力较弱;任务与配置间鲁棒性差;效率约束普遍存在;毒性与偏见问题报告不足。主要挑战集中在经济可行性、评估有效性与社会技术融合。未来方向建议提升领域感知模型能力,建立全面、标准化的评估体系。结论:基于大模型的代码生成任务发展迅速但评估不均,亟需提升领域感知能力与建立统一、全面的评估机制,同时关注效率与成本问题。

原文摘要 · Abstract (English)

Context. Large language models (LLMs) are increasingly applied to code-generating tasks (CGTs) in software engineering. While reported results are promising, the broader effects of such application and their integration into real-world development remain insufficiently understood with existing tertiary studies provide little in this area. Objective. This tertiary study consolidates secondary evidence on LLM-based CGTs, synthesizing the publication landscape, effects, scenarios, integration challenges, and future research directions. Method. Following systematic review guidelines, we searched in related digital libraries, complemented by backward-and-forward snowballing and screening step. Study quality was assessed and extraction reliability was audited with inter-rater agreement statistics. Evidence was synthesized using SWEBOK knowledge areas and the HELM framework. Results. We identify 30 secondary studies published between 2017-2025, with rapid growth since 2023. Accuracy seems strong on benchmarks but weakly supported for real-world generalization; robustness is fragile across tasks and configurations; efficiency constraints are pervasive; toxicity and bias are under-reported. Dominant challenges concern economic feasibility, evaluation validity, and socio-technical integration. Future directions suggest domain-aware model improvement and the need for holistic, standardized evaluation. Conclusion. LLM-based CGTs represent a fast-maturing yet unevenly evaluated research area, highlighting the need for domain-aware model improvements and holistic, standardized evaluation, addressing efficiency and associated costs.

代码生成大模型综述评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。