AI可全自动修复代码并生成需求,准确率超人类专家。
From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization
- 用大模型结合形式验证与测试驱动,迭代优化代码修复
- 在SWE-bench上达48.33%准确率,比前人提升38.6%
- 适合追求自动化开发的工程团队和研究者
本文标志着人工智能与软件工程融合的新纪元,将机器置于编程能力的顶峰。我们提出一种形式化、迭代的方法,证明了AI可在代码创建与优化的所有环节替代人类程序员。该方法结合大语言模型、形式化验证、测试驱动开发及增量架构指导,在SWE-bench基准上实现48.33%的准确率,较当前最优表现提升38.6%。这一成果突破了此前认为的上限,宣告了人类独占编码时代的终结,并开启了由AI驱动的自主软件创新新时代。本工作不仅是一项技术进步,更挑战了关于人类创造力持续数个世纪的假设。我们提供了强有力的证据,证明AI在实际工程场景中具有优越性,为计算创造力超越人类才智的未来奠定了基础。
原文摘要 · Abstract (English)
This manuscript signals a new era in the integration of artificial intelligence with software engineering, placing machines at the pinnacle of coding capability. We present a formalized, iterative methodology proving that AI can fully replace human programmers in all aspects of code creation and refinement. Our approach, combining large language models with formal verification, test-driven development, and incremental architectural guidance, achieves a 38.6% improvement over the current top performer's 48.33% accuracy on the SWE-bench benchmark. This surpasses previously assumed limits, signaling the end of human-exclusive coding and the rise of autonomous AI-driven software innovation. More than a technical advance, our work challenges centuries-old assumptions about human creativity. We provide robust evidence of AI superiority, demonstrating tangible gains in practical engineering contexts and laying the foundation for a future in which computational creativity outpaces human ingenuity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。