arXiv:2502.00226cs.LGcs.SE2025-02

评测大模型在跨领域多文件项目中的代码正确性与一致性

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems

  • 设计真实项目场景的多文件编程任务,模拟实际开发
  • 32次运行显示顶级模型平均得分75%,Claude-3.5表现最稳定
  • 适合关注模型可靠性与工程落地能力的研究者

评估大语言模型(LLMs)在真实世界软件开发任务中的适用性具有重要意义。现有基准大多聚焦于单文件编码问题或特定库,忽视了多文件项目场景,且缺乏对一致性的严格评估。HackerRank-ASTRA 基准引入了模拟真实场景的项目式编程问题,通过32次运行(k=32)和中位数标准差评估模型一致性,并结合层级分类法分析子技能表现。在65个问题上的初步评估显示,top三模型——o1、o1-preview 和 Claude-3.5-Sonnet-1022——平均得分均达75%,性能无显著差异。值得注意的是,Claude-3.5-Sonnet-1022 在所有问题上表现出最高一致性,变异系数低(SD = 0.0497),统计上显著优于其他模型,凸显其在真实开发任务中的可靠性。

原文摘要 · Abstract (English)

Evaluating the real-world applicability of large language models (LLMs) provides valuable insights for their development and use in software development tasks. Existing benchmarks often focus on standalone coding problems or specific libraries, overlooking multi-file, project-based scenarios and lacking a rigorous evaluation of consistency. The HackerRank-ASTRA Benchmark introduces project-based coding problems that mirror real-world scenarios. It evaluates model consistency through 32 runs (k = 32) and median standard deviation while incorporating taxonomy-level analysis to assess sub-skill capabilities. Initial evaluations on 65 problems show that the top three models -- o1, o1-preview, and Claude-3.5-Sonnet-1022 -- achieved comparable average scores of 75%, with no statistically significant differences in performance. Notably, Claude-3.5-Sonnet-1022 demonstrated the highest consistency across problems, with low variability (SD = 0.0497), which was statistically significant compared to other models, highlighting its reliability for real-world software development tasks.

大模型评估代码生成一致性评测真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。