arXiv:2503.07010cs.SEcs.CL2025-03ACL被引 27

构建用户视角的代码生成评估基准,提升编程智能体实用性。

ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation

  • 通过模拟用户交互自动评估项目级代码生成效果。
  • 发现系统工程理解与综合分析是生成实用代码的关键。
  • 适合研究编程智能体评估与真实场景应用的研究者。

近期,大语言模型(LLM)智能体在编程能力上取得快速进展。然而,现有基准难以从用户视角自动评估,且结果缺乏可解释性。为此,我们提出 ProjectEval,一个基于大模型生成并经人工审核的项目级代码生成自动化评估基准。该基准包含自然语言或代码骨架三种不同层级输入,支持通过用户交互模拟进行执行评估,并结合现有客观指标计算代码相似度。实验表明,系统化工程理解、项目整体把握与综合分析能力是 LLM 智能体实现实际项目的核心因素。本研究的发现与基准为开发更高效、可部署于真实生产环境的编程智能体提供了重要参考。

原文摘要 · Abstract (English)

Recently, LLM agents have made rapid progress in improving their programming capabilities. However, existing benchmarks lack the ability to automatically evaluate from users' perspective, and also lack the explainability of the results of LLM agents' code generation capabilities. Thus, we introduce ProjectEval, a new benchmark for LLM agents project-level code generation's automated evaluation by simulating user interaction. ProjectEval is constructed by LLM with human reviewing. It has three different level inputs of natural languages or code skeletons. ProjectEval can evaluate the generated projects by user interaction simulation for execution, and by code similarity through existing objective indicators. Through ProjectEval, we find that systematic engineering project code, overall understanding of the project and comprehensive analysis capability are the keys for LLM agents to achieve practical projects. Our findings and benchmark provide valuable insights for developing more effective programming agents that can be deployed in future real-world production.

编程智能体代码生成自动化评估项目级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。