AI写代码让评审变快但争议大,研究揭示人为因素决定其影响方向。
3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
- 从3100篇工程师讨论中构建因果理论,揭示评审机制如何受团队控制
- 发现AI提价合并更快但审查更少,结果依赖分析方式而波动
- 提供可复用的灰文献建模方法,适合关注AI与工程实践的研究者
当前编码代理能自动生成完整拉取请求,从业者对其对代码评审的影响存在激烈分歧:是否成为瓶颈、人工评审是否仍必要、是否悄然削弱理解力。现有仓库挖掘研究仅捕捉表面趋势,难解释内在机制,且趋势本身不稳定。对公开GitHub活动的观察分析显示,代理生成的拉取请求被审查频率更低、合并速度快三倍以上、讨论也更少,但不同合理分析策略下趋势方向反转,仅揭示变化现象而未说明原因。为还原机制,本研究大规模合成从业者话语:收集38,709份灰文献(工程博客与Reddit帖文),筛选出3,100篇实质性讨论代码评审的内容,通过LLM辅助管道编码,构建包含26个构念与67条关系(64条有向,3条争议)的因果模型。核心观点是:评审是决定编码代理对软件影响的关键控制点,AI不固定其效应方向——最终由团队的人才能力与评审流程设计决定。该理论使对立立场显性化,并将‘AI改变评审’转化为可检验的命题。次要贡献为提供可扩展的灰文献理论构建方法,附公开实现。
原文摘要 · Abstract (English)
Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes the bottleneck, whether human review is still necessary, and whether it quietly erodes the understanding that it once built. Repository-mining studies measure surface trends but seldom explain the mechanisms beneath them, and the trends themselves prove unstable. A motivating observational analysis of public GITHUB activity finds that agent-authored pull requests are reviewed less often, merged several times faster, and discussed less than human-authored ones, yet the direction of these trends flips under different but equally defensible analysis choices, so the traces establish what is changing without explaining why. To recover the mechanisms, we synthesize practitioner discourse at scale into an explanatory theory: we collect 38,709 grey-literature documents (engineering blogs and Reddit threads), filter to those substantively about code review, and code a stratified random sample of 3,100 with an LLM-assisted pipeline, from which we build a causal model of 26 constructs and 67 relationships (64 directed, 3 contested). Its organizing claim is that review is the control point through which a coding agent's effect on software is decided, and that AI does not fix the sign of that effect: the team sets it, through the expertise its humans bring and how it structures the review process. The theory makes the competing positions explicit and turns "AI is changing code review" into falsifiable propositions with named constructs and moderators. As a secondary contribution, we offer the underlying LLM-assisted, grey-literature theory-building method as a scalable template for software-engineering research, with a public implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。