构建全面的代码大模型评估框架,解决现有基准测试的局限性。
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
- 提出BEHELM框架,整合软件场景与多维度评估
- 揭示三大评估障碍:数据匮乏、指标单一、流程不统一
- 适合研究者和开发者用于公平、真实地评测代码生成模型
代码领域的大语言模型发展迅速,但评估能力滞后。现有基准测试聚焦狭窄任务和单一指标,掩盖了鲁棒性、可解释性、公平性、效率和实际可用性方面的关键缺陷。它们还存在数据工程不一致、软件工程上下文不足及广泛的数据污染问题。我们通过深入调研现有基准并结合专题社区研讨会的见解,识别出三大核心障碍:缺乏丰富的软件工程数据集、过度依赖机器学习中心指标、缺乏标准化可复现的数据流水线。基于此,我们提出BEHELM,一个融合软件场景定义与多指标评估的综合性基准框架。该框架结构化地评估模型在不同任务、编程语言、输入输出粒度及关键质量维度的表现。目标是降低构建基准的开销,实现对代码大模型在软件工程中应用的公平、真实且面向未来的评估。
原文摘要 · Abstract (English)
Large language models for code are advancing fast, yet our ability to evaluate them lags behind. Current benchmarks focus on narrow tasks and single metrics, which hide critical gaps in robustness, interpretability, fairness, efficiency, and real-world usability. They also suffer from inconsistent data engineering practices, limited software engineering context, and widespread contamination issues. To understand these problems and chart a path forward, we combined an in-depth survey of existing benchmarks with insights gathered from a dedicated community workshop. We identified three core barriers to reliable evaluation: the absence of software-engineering-rich datasets, overreliance on ML-centric metrics, and the lack of standardized, reproducible data pipelines. Building on these findings, we introduce BEHELM, a holistic benchmarking infrastructure that unifies software-scenario specification with multi-metric evaluation. BEHELM provides a structured way to assess models across tasks, languages, input and output granularities, and key quality dimensions. Our goal is to reduce the overhead currently required to construct benchmarks while enabling a fair, realistic, and future-proof assessment of LLMs in software engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。