arXiv:2409.18952cs.SEcs.LG2024-09被引 28

构建首个面向前沿代码修复模型的标准化评估基准

RepairBench: Leaderboard of Frontier Models for Program Repair

  • 基于实际编译执行测试的评估机制,确保修复结果真实有效
  • 在Defects4J和GitBug-Java数据集上持续评测前沿模型表现
  • 公开评估框架,支持新模型动态更新榜单,助力研究进展追踪

AI驱动的程序修复利用AI模型生成补丁来修复有缺陷的软件。随着AI技术的快速进步,程序修复的最先进性能不断变化,但要把握这一进展,需要频繁且标准化的评估。我们提出RepairBench,一个面向AI驱动程序修复的新基准排行榜。其关键特性包括:1)基于执行的评估:所有补丁均经过编译并在测试套件上执行;2)对前沿模型进行定期、标准化评估。RepairBench采用两个高质量基准Defects4J和GitBug-Java,评估前沿模型在真实世界程序修复任务中的表现。我们公开发布RepairBench的评估框架,并将持续更新榜单以纳入新发布的前沿模型。

原文摘要 · Abstract (English)

AI-driven program repair uses AI models to repair buggy software by producing patches. Rapid advancements in AI surely impact state-of-the-art performance of program repair. Yet, grasping this progress requires frequent and standardized evaluations. We propose RepairBench, a novel leaderboard for AI-driven program repair. The key characteristics of RepairBench are: 1) it is execution-based: all patches are compiled and executed against a test suite, 2) it assesses frontier models in a frequent and standardized way. RepairBench leverages two high-quality benchmarks, Defects4J and GitBug-Java, to evaluate frontier models against real-world program repair tasks. We publicly release the evaluation framework of RepairBench. We will update the leaderboard as new frontier models are released.

程序修复AI评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。