arXiv:2606.20882cs.SEcs.AI2026-06

AI生成代码让作者归属无法反映真实理解,旧度量方法已失效。

The Substrate Collapse: AI Code Generation Invalidates Authorship-Based Knowledge Metrics

  • 用人类作者身份推断知识掌握程度的逻辑被AI生成代码打破。
  • 作者足迹仍存在,但不再能证明理解程度,指标失去意义。
  • 未来需基于理解证据构建新度量,适合关注系统可信度的研究者。

软件工程长期通过代码作者来推断系统知识的归属。卡车因子、作者度量和知识度模型均基于一个假设:编写某段代码即意味着理解它。历史上这一假设成立,因代码入库前必须由人类撰写,至少带来临时理解。本文指出,AI代码生成从根本上切断了这一推断链,导致以作者为基础的知识度量不仅退化,更整体失效。当智能体生成模块而人类仅合并时,版本记录仍归功于该作者,但此归属不再支持理解结论——相同作者痕迹可能对应完全、部分或无理解。度量仍返回数值,但其衡量的底层实体已与原目标脱钩。该崩溃得到领域自身测量失败的佐证。方法论上,无法通过优化作者足迹函数恢复原有推断,因为足迹不再支撑该推理。替代方案必须建立在理解证据之上。本文提出可验证预测:具有健康作者衍生卡车因子但低理解留存率的系统,将出现作者度量无法预判的故障修复失败。构建系统级与团队级的理解导向度量,是当前领域核心未解测量难题。

原文摘要 · Abstract (English)

Software engineering has long inferred where a system's knowledge resides from who authored its code. The truck factor, the Degree-of-Authorship metric, and the degree-of-knowledge model all rest on one inference -- that authoring a region of code is evidence of understanding it -- and for most of software's history it was a workable proxy, because code entered a repository only when a human wrote it, which forced at least transient understanding. This paper argues that AI code generation severs that inference at its root, and that the consequence is not the degradation of the authorship-based metrics but their invalidation as a class. When an agent generates a module and a human merges it, the version-control record still attributes authorship, but the attribution no longer licenses any conclusion about comprehension: the same footprint is now compatible with full, partial, or no understanding. The metric still returns a number; the number measures a substrate that has come uncoupled from the quantity it was used to estimate. The collapse is corroborated by the field's own measurement failures, and the methodological corollary is load-bearing: the instrument the comprehension-debt era needs cannot be built by refining the knowledge-concentration metrics, because no function of an authorship footprint recovers an inference the footprint no longer supports. The replacement must be grounded in evidence of comprehension rather than authorship. I state a falsifiable prediction that discriminates the two -- that systems with a healthy authorship-derived truck factor but low comprehension-measured retention will suffer incident-resolution failures the authorship metric does not predict -- and argue that building the comprehension-grounded instrument at the scale of a system and a team is the field's open measurement problem, left open here.

AI代码度量失效知识归属理解验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。