arXiv:2607.01087cs.SEcs.AI2026-07被引 3

用12周实证研究揭示AI编程如何从失控到可管的转化机制

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

  • 通过真实开发记录构建治理转化模型,揭示失败如何催生管控机制
  • 产出420KLOC生产代码与1.16MLOC测试/工具代码,验证高并发生成的可管理性
  • 适合关注AI辅助开发治理、工程实践创新的研究者与工程师

生成式AI正将软件工程从依赖稀缺编码能力转向高产低耗的代码生成模式。这一转变使核心问题从‘AI能否生成有用代码’转变为‘如何组织架构、工具、证据和反馈环,以确保AI辅助开发过程可检查、可修正、可持续维护’。本文通过第一人称案例研究,考察一名专家工程师在12周内使用前沿AI编程代理构建文档可访问性修复系统的过程。实证数据包括88份实时工作笔记、420千行生产代码,以及116万行测试、静态检查、支持文档和代理工具代码。基于此,我们提出一个关于治理转换的中层理论,以过程模型解释高速代理式开发如何演化为可治理状态:代理开发速度会暴露反复出现的结构性失败,而工程判断则通过将这些失败转化为持久的治理机制来维持开发速度。与传统从既定责任推导控制的模型不同,治理转换强调控制是在代理工作中浮现的失败中被发现的。我们利用该模型做出可验证预测,并探讨其对软件工程研究与实践的影响。

原文摘要 · Abstract (English)

Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI can generate useful code, but how engineers organize architectures, tools, evidence, and feedback loops so that AI-mediated development remains inspectable, correctable, and maintainable. We study this problem through a first-person case study: a 12-week development effort in which a single expert software engineer used frontier AI coding agents to build a document accessibility remediation system. The empirical record comprises 88 contemporaneous field notes, 420 KLOC of production code, and 1.16 MLOC of tests, lints, supporting documentation, and agent tooling. From this record, we develop a candidate middle-range theory of governance conversion, expressed as a process model explaining how high-velocity agentic implementation becomes governable. The model explains how agentic implementation velocity surfaces recurring structural failure classes, and how engineering judgment sustains velocity by converting those failures into durable governance mechanisms. In contrast to existing governance models that derive controls from known obligations, governance conversion explains how controls are discovered from failures that become visible only during agentic work. We use our model to make testable predictions and to describe implications for software engineering research and practice.

AI编程治理机制工程实践

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。