评测大模型代码生成能力,发现评估结果可能被高估。
An evaluation of LLM code generation capabilities through graded exercises
- 用8种语言的分级编程题测试GPT4-o-mini模型
- 成功率随题目难度、语言流行度和发布时间增长而提高
- 近四成表现源于题目泄露,非真实编码能力
大型语言模型在从自然语言生成功能性代码方面展现出显著能力,但目前尚缺乏客观、无偏的标准化评估方法。本文综述现有评估手段,并对当前最先进模型GPT4-o-mini在Codewars平台获取的8种编程语言的精选编程挑战中的表现进行新评估。分析显示,模型成功概率与任务难度、所用编程语言的流行度以及挑战发布以来的时间呈正相关。进一步的高阶特征解释分析表明,约46.6%的模型表现可归因于任务难度,37.4%可能源于挑战答案在训练数据中的泄露,剩余16%则与编程语言有关。这些结果提示,现有评估方法可能过高估计了大模型生成功能性代码的真实能力。
原文摘要 · Abstract (English)
Large Language Models have shown prominent capabilities in generating functional code from natural language descriptions. However, a standardized way to evaluate these capabilities in an objective and unbiased manner is still to be found. In this paper we review the current evaluation methods available to this end, and run a new evaluation of the performance of one state-of-the-art model (GPT4-o-mini) in solving curated coding challenges in 8 programming languages, obtained from Codewars, a software development community. Our analysis shows that the chance of success of the model has a positive correlation with the task difficulty, the popularity of the programming language being used and the time elapsed since the publication of the challenge. A further approximate explanatory analysis in terms of high-level features hints that while 46.6% of the model performance could be attributed to task difficulty, a 37.4% seems to be related to leakage of the challenge solutions into the model training set, while the remaining 16% depends on the programming language. These results suggest that current evaluation methodologies might be overestimating the actual skill of Large Language Models for generating functional code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。