arXiv:2501.11354cs.SEcs.AI2025-01被引 9

构建六层框架,系统梳理大模型代码生成的挑战与路径

Towards Advancing Code Generation with Large Language Models: A Research Roadmap

  • 提出六层分阶段框架,覆盖从输入到验证的全流程
  • 分析大模型在真实开发中面临的可靠性与评估难题
  • 适合研究代码生成的开发者与系统设计者参考

近年来,大型语言模型在代码生成任务中展现出卓越能力,但其在实际开发场景中的应用仍面临诸多技术和评估挑战。本文提出一种六层视觉框架,将代码生成过程划分为输入、编排、开发和验证等阶段,并深入分析现有研究与主流框架。针对基于大模型的智能体系统在代码生成中遇到的问题,系统性地梳理了当前技术瓶颈。在此基础上,本文提供了多维度视角与可操作建议,旨在提升大模型代码生成系统的可靠性、鲁棒性和可用性。本工作致力于解决长期存在的难题,为未来更实用的大模型代码生成方案提供实践指导。

原文摘要 · Abstract (English)

Recently, we have witnessed the rapid development of large language models, which have demonstrated excellent capabilities in the downstream task of code generation. However, despite their potential, LLM-based code generation still faces numerous technical and evaluation challenges, particularly when embedded in real-world development. In this paper, we present our vision for current research directions, and provide an in-depth analysis of existing studies on this task. We propose a six-layer vision framework that categorizes code generation process into distinct phases, namely Input Phase, Orchestration Phase, Development Phase, and Validation Phase. Additionally, we outline our vision workflow, which reflects on the currently prevalent frameworks. We systematically analyse the challenges faced by large language models, including those LLM-based agent frameworks, in code generation tasks. With these, we offer various perspectives and actionable recommendations in this area. Our aim is to provide guidelines for improving the reliability, robustness and usability of LLM-based code generation systems. Ultimately, this work seeks to address persistent challenges and to provide practical suggestions for a more pragmatic LLM-based solution for future code generation endeavors.

代码生成LLM研发框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。