arXiv:2511.08301cs.AIcs.SE2025-11被引 4

让AI编程助手像人类开发者一样共同学习,提升代码质量。

Smarter Together: Creating Agentic Communities of Practice through Shared Experiential Learning

  • 构建共享记忆库Spark,让AI代理可共同积累和使用经验。
  • 小模型(300亿参数)在Spark帮助下达到大模型水平的代码质量。
  • 推荐结果在最佳实践上帮助率达98.2%,适合协作式AI开发场景。

从以人为主转向以智能体为主软件开发实践,正在颠覆现有开发者知识共享环境。传统开发者社区的知识分享参与度在短时间内大幅下降,而能替代人类的智能体功能尚未出现,导致已生成大量新代码的AI代理无法获取有价值的集体经验。本文提出Spark——一种模拟人类开发者社区集体智慧的共享智能体记忆架构。Spark使同一问题领域的AI编码代理能持续贡献并获取经验,实现群体持续学习。我们评估Spark作为AI编码代理的教练效果,发现其建议显著提升不同规模通用代码生成模型的代码质量。借助Spark,一个300亿参数的小型开源模型达到大型顶尖模型的代码质量水平。此外,我们基于软件开发最佳实践标准评估Spark推荐的内在质量,在五档中前两档的定性帮助率最高达98.2%。

原文摘要 · Abstract (English)

The transition from human-centric to agent-centric software development practices is disrupting existing knowledge sharing environments for software developers. Traditional peer-to-peer repositories and developer communities for shared technical knowledge and best practice have witnessed dramatic drops in participation in a short period of time. At the same time, agentic functional equivalents are yet to emerge leaving AI agents, which already generate a significant proportion of all new software code produced, without access to repositories of valuable shared learning. In this paper, we introduce Spark, a novel shared agentic memory architecture which is designed to emulate the collective intelligence and know-how of human developer communities. Spark enables AI coding agents to both contribute to and draw from a persistent and continuously evolving experiential memory. Agents operating in the same general problem space use the Spark shared memory as a repository of new knowledge to achieve collective continual learning. We evaluate Spark as a coach for AI coding agents performing software development tasks. We demonstrate that recommendations made by Spark improve the quality of code generated by generic code generation models at varying sizes and capability tiers. Boosted by Spark, a small open-weights model with 30 billion parameters was able to match the code quality afforded by a much larger state-of-the-art model. Separately, we measure the intrinsic quality of recommendations generated by Spark against a wide range of criteria inspired by software development best practice, and achieve helpfulness levels of up to 98.2% in the top two (out of five) qualitative helpfulness bands.

AI编程集体智能代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。