为大模型代码生成研究提供标准化评估框架
Designing Empirical Studies on LLM-Based Code Generation: Towards a Reference Framework
- 构建涵盖问题来源、质量属性和度量指标的评估框架
- 通过案例映射验证框架可操作性,提升研究可比性
- 适合从事大模型代码生成实验设计的研究者参考
大语言模型在自动化代码生成方面展现出变革性潜力,可应对广泛的软件工程挑战。然而,当前基于大模型的代码生成实证研究缺乏统一标准,研究目标、任务和度量指标差异较大,限制了结果的可比性和可复现性。本文提出一个理论框架,用于设计和报告大模型代码生成的实证研究。该框架基于我们过往的实验经验,并通过对近期关键研究的比较分析,系统组织评估的核心要素,包括问题来源、质量属性和度量指标,支持结构化与系统化的实验设计。我们通过代表性案例映射展示了其适用性,并识别出优化空间。未来计划将该框架发展为更稳健成熟的工具,以推动大模型评估在软件工程领域的标准化。
原文摘要 · Abstract (English)
The rise of large language models (LLMs) has introduced transformative potential in automated code generation, addressing a wide range of software engineering challenges. However, empirical evaluation of LLM-based code generation lacks standardization, with studies varying widely in goals, tasks, and metrics, which limits comparability and reproducibility. In this paper, we propose a theoretical framework for designing and reporting empirical studies on LLM-based code generation. The framework is grounded in both our prior experience conducting such experiments and a comparative analysis of key similarities and differences among recent studies. It organizes evaluation around core components such as problem sources, quality attributes, and metrics, supporting structured and systematic experimentation. We demonstrate its applicability through representative case mappings and identify opportunities for refinement. Looking forward, we plan to evolve the framework into a more robust and mature tool for standardizing LLM evaluation across software engineering contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。