arXiv:2412.05288cs.SEcs.CL2024-12NeurIPS被引 22

构建两个新基准,评估大模型在编程协助中的真实表现。

StackEval: Benchmarking LLMs in Coding Assistance

  • 基于Stack Overflow数据构建大规模评测集
  • 发现大模型对新内容处理能力有限
  • 适合研究者与开发者评估模型编程能力

我们提出了两个全面的基准,用于评估语言模型在编程协助任务中的表现,涵盖代码生成、调试、代码审查和概念理解。主要贡献包括两个精心构建的数据集:StackEval,一个基于Stack Overflow问题的大规模基准;以及StackUnseen,一个包含最新内容的动态基准。这些基准为理解大模型在处理新兴内容时的能力与局限提供了新视角。此外,我们还利用人工标注数据集评估了大模型作为代码任务评判者的性能,探索其评估能力及潜在偏见,包括是否偏好自身生成的解决方案。研究结果凸显了这些基准在推动大模型在编程协助领域发展与应用中的潜力。为确保可复现性,我们已将数据集和评估代码公开于https://github.com/ProsusAI/stack-eval。

原文摘要 · Abstract (English)

We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack Overflow questions, and StackUnseen, a dynamic benchmark featuring the most recent Stack Overflow content. These benchmarks offer novel insights into the capabilities and limitations of LLMs, particularly in handling new and emerging content. Additionally, we assess LLMs' proficiency as judges for coding tasks using a curated, human-annotated dataset, exploring their evaluation capabilities and potential biases, including whether they favor their own generated solutions. Our findings underscore the potential of these benchmarks to advance LLM development and application in coding assistance. To ensure reproducibility, we publicly share our datasets and evaluation code at https://github.com/ProsusAI/stack-eval .

大模型编程评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。