arXiv:2410.04485cs.SEcs.AI2024-10被引 2

用对话式方法修复代码,成功率近一半,媲美顶尖水平。

Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • 基于LLaMA 3.1 70B构建简单对话流水线生成补丁。
  • 在SWE-Bench上成功生成有效补丁率达47%。
  • 为自动化修复提供新思路,适合研究代码生成与调试的开发者。

项目级自动程序修复可能在人类活动多个领域开启新机遇。自SWE-Bench挑战提出以来,已有众多解决方案涌现。补丁生成是程序修复的重要环节,基于测试套件的对话式补丁生成已证明其有效性。然而,对话式补丁生成在SWE-Bench上的潜力尚未被专门评估。本研究报告了针对SWE-Bench问题的对话式补丁生成独立有效性的实验结果。实验表明,基于LLaMA 3.1 70B的简单对话流程可在47%的情况下生成有效补丁,该表现与当前SWE-Bench上程序修复的最先进水平相当。

原文摘要 · Abstract (English)

Automatic program repair at project level may open yet to be seen opportunities in various fields of human activity. Since the SWE-Bench challenge was presented, we have seen numerous of solutions. Patch generation is a part of program repair, and test suite-based conversational patch generation has proven its effectiveness. However, the potential of conversational patch generation has not yet specifically estimated on SWE-Bench. This study reports experimental results aimed at evaluating the individual effectiveness of conversational patch generation on problems from SWE-Bench. The experiments show that a simple conversational pipeline based on LLaMA 3.1 70B can generate valid patches in 47\% of cases, which is comparable to the state-of-the-art in program repair on SWE-Bench.

程序修复对话生成LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。