代码基准需重视严谨性、可靠性和可复现性
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
- 提出55项检查清单,规范代码基准构建流程
- 2025年忽略代码覆盖率的基准数等于过去十年总和
- 适合评估大模型能力的研究者与开发者参考
代码相关基准在评估大语言模型(LLMs)中起关键作用,但其质量直接影响社区对模型能力的理解。尽管近年来对基准质量的关注度上升,但我们对2014至2025年间672个代码基准的跨十年调研发现,认知与实践仍存在显著滞后。例如仅2025年,不提供代码覆盖率测试用例的基准数量几乎等同于此前十年累计总数。为此,我们明确主张:代码基准必须优先保障构建严谨性、评估可靠性与发布可复现性。为实现这一目标,我们提出名为HOW2BENCH的代码基准指南,包含55项检查清单。此外,人工研究进一步揭示,当前问题不仅源于实施难度大,更因对重要性认识不足。
原文摘要 · Abstract (English)
Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In the past few years, awareness of benchmark quality has grown. Yet, after a decade-scale (2014-2025) survey over 672 code benchmarks, we observed a lag between growing awareness and actual practice. For example, in 2025 alone, the number of benchmarks that ignore code coverage when providing test cases nearly matches the total count accumulated across the previous ten years. In response, we take a clear position: Code benchmarks must prioritize rigor in benchmark construction, reliability in evaluation, and reproducibility in release. To operationalize this position, we introduce a code benchmark guideline HOW2BENCH with 55 checklists. Finally, our further human study also exposed that the current issues not only stem from the significant effort required, but also from a lack of awareness regarding their importance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。