AI编码助手提升效率,但削弱了开源社区的公共知识积累。
From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

- 用真实数据模拟开发者与AI代理并行协作,对比有无代理的效果。
- 引入代理后任务完成率升39%,但仅26%开发者使用,且集中在活跃者。
- 公开知识质量下降,后续开发者检索难度增加,影响社区可持续性。
开源软件社区是数字公共基础设施,不仅生成代码,还通过可见协作创造公共知识和人际关系。生成式编码代理(CAs)可提升开发效率,但将部分活动从公开的人类互动转向私密的人机循环。本研究基于1,084名活跃开发者的实际GitHub数据,构建基于LLM的多代理仿真系统,经历史提交预热后,将同一社区状态分为无代理(No-CA)与有代理(CA)两组进行4周模拟。引入CA后,计划与完成的任务量分别提升34.0%和39.0%,中位完成时间从45分钟降至20分钟。然而,代理采纳率仅为26.0%,增益集中于原本更活跃、连接度更高的开发者。任务执行路径也被重构:人与人直接互动占比从32.4%降至11.6%,而涉及代理的模式升至57.3%,其中40.3%由代理辅助的自我循环完成。在标准检索基准上,CA生成的知识库覆盖率仅22.3%,远低于真实人类贡献的81.1%,且需更多检索步骤、成功率更低。结果揭示出生产力与公共知识间的张力:编码代理虽提升技术产出,但更多工作转入代理中介或私密循环,使公共记录对后续贡献者支持减弱。
原文摘要 · Abstract (English)
Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced tool to improve development efficiency while shifting part of activities from public human interaction to private human-agent loops. We study this shift using an LLM-based multi-agent simulation initialized with real GitHub data from 1,084 active developers and their repository relationships. After a warm-up with historical commits, we branch the same community state into parallel No-CA and CA conditions for 4-week simulations. CA introduction increases planned and completed tasks by 34.0% and 39.0%, respectively, and reduces median completion time from 45 to 20 minutes. However, adoption reaches only 26.0%, and the gains concentrate among developers who are already more active and well connected. CAs also restructure task execution pathways. Direct human-human interaction declines from 32.4% to 11.6%, while CA-involved modes increase to 57.3%, including 40.3% completed through CA-assisted self-loops. Public knowledge generated under CA condition also provides less support for later tasks. On a standardized retrieval benchmark, the CA corpus achieves 22.3% knowledge coverage, far below the 81.1% achieved by the real-human corpus, and requires more retrieval steps with a lower success rate. These results reveal a productivity-public knowledge tension: coding agents increase technical production, but more work shifts to agent-mediated or private loops, leaving public records less useful to future contributors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。