分析智能代码工具合并后维护负担,发现其更易出漏洞且依赖风险更高。
Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

- 对比182个仓库中人工与智能代码的长期维护情况
- 智能代码修正率更高,安全缺陷和依赖漏洞更多
- 项目无审查率每增10%,智能代码负担平均升6%
智能编码工具越来越多地用于自主修改真实项目。现有研究主要关注提交前的接受率与评审工作量,对合并后的结果关注甚少。然而,合并成功并不意味着代码能长期稳定。我们对182个仓库中的智能与人工贡献进行了纵向实证分析,追踪其合并后命运,分析后续修改意图及引入的缺陷与漏洞。结果显示,尽管总体维护率相似,但智能代码需更高修正频率,引入更多安全弱点与依赖漏洞。统计显著表明,项目无审查率每上升10个百分点,智能代码的维护负担平均增加约6%。随着编码代理在开发中普及,必须不仅评估其能否通过合并,更要确保其长期安全与可维护性。
原文摘要 · Abstract (English)
Agentic coding tools are increasingly used to make autonomous repository-level changes to real-world projects. Prior work has largely evaluated these contributions at the pre-merge stage, through outcomes such as pull request acceptance and review effort. Far less is known about what happens to agentic code post-merge. Yet merge success alone does not reveal whether a contribution will remain stable or require bug fixes and other corrective maintenance downstream. We conduct a longitudinal empirical analysis of agentic and human contributions across 182 repositories, tracking their post-merge fate over time, characterizing the intent of subsequent modifications, and analyzing the defects and vulnerabilities they introduce. While the overall maintenance rates are similar, agentic contributions require significantly higher rates of corrective maintenance and introduce more security weaknesses and dependency vulnerabilities. We also find statistically significant evidence that agentic maintenance burden is associated with repository characteristics. In particular, each 10 percentage-point increase in a project's no-review rate is associated with roughly a 6% increase in agentic maintenance burden on average. As coding agents become pervasive in software development, our findings highlight the need to evaluate and design agentic tools not only to produce mergeable changes, but to produce contributions that remain secure and maintainable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。