arXiv:2606.22711cs.SEcs.AI2026-06中稿 · KDD

AI协作提代码更难通过,但这是假象,背后有层层数据陷阱。

Beyond Simpson's Paradox: A Cascade of Confounders in AI Agent Pull-Request Co-Authorship

  • 按不同AI工具分层分析,发现多数工具的协作提审反而更易通过
  • 控制仓库和提交次数后,协作优势基本消失,说明是数据偏差所致
  • 适合关注数据偏见、评估AI工具真实效果的研究者阅读

综合五款AI编程工具的数据,带人类合作者标记的代码提交(PR)合并率(53.8%)低于纯自主提交(79.8%),但这属于典型的辛普森悖论。对来自AIDev数据集的33,596个PR按工具分层后发现:Copilot与Devin在各自内部显示显著正向差距(+41.2和+33.5个百分点,均p<0.001),而Cursor、Claude Code和Codex则效应微弱且置信区间包含零。该悖论完全由工具分布驱动:Codex占数据集64.9%,合并率高但极少使用合作者标记。然而,辛普森悖论只是多重混杂因素的第一层:控制项目后,Devin的优势从+33.5降至+1.6(p=0.73);进一步控制提交次数,Copilot的组内差距从+36.2降至+24.4;仅限多提交的PR时,其效果降为+4.8(p=0.59)。一旦同时控制仓库选择与PR结构,各工具均无明显协作优势。研究警示:报告聚合数据时必须分层,跨段落的协作关联实为选择偏差与提交结构伪影,而非因果效益。

原文摘要 · Abstract (English)

Pooled across five AI coding agents, pull requests (PRs) with a human Co-Authored-By trailer merge less often than purely-autonomous ones (53.8% vs. 79.8%) -- yet this aggregate finding is a textbook Simpson's Paradox. Stratifying 33,596 PRs from the AIDev dataset by agent identity reverses the conclusion: Copilot and Devin show large positive within-agent gaps (+41.2 and +33.5 pp, both p<0.001), while Cursor, Claude Code, and Codex show small effects whose cross-sectional 95% CIs span zero. The paradox is driven entirely by agent composition: Codex, which dominates 64.9% of the dataset, achieves high merge rates while rarely using co-authorship. But Simpson's Paradox is only the first layer of a cascade of confounders: within-repo controls eliminate Devin's gap (+33.5 to +1.6 pp, p=0.73); a commit-count control further halves Copilot's within-repo gap (+36.2 to +24.4 pp); restricted to multi-commit PRs, the Copilot within-repo effect dissolves to +4.8 pp (p=0.59). No agent retains a clear co-authorship effect once both repository selection and PR structure are controlled. Our findings caution against reporting agent-pooled statistics without stratification and demonstrate that cross-sectional co-authorship associations are largely selection and PR-structure artefacts rather than evidence of a causal benefit.

AI编程数据偏见因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。