arXiv:2604.16323cs.SEcs.AI2026-04中稿 · the Human-Centered…

提出新框架,追踪智能编程代理的决策演化,防止意图漂移。

Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development

  • 通过时间、工具调用和架构边界追踪代理行为,构建过程解释框架。
  • 在两类用户中验证:普通用户获结构可见性,专业开发者提升审代码上下文。
  • 引入因果图界面与意图级监控,实现轻量级人机协同监督。

随着自主编程代理深度嵌入软件开发流程,其高运作速度带来了关键的监管盲区:代理行为与架构意图之间的持续偏离。我们称这一现象为‘代理熵’——传统基于代码diff或HCXAI的方法无法捕捉,因其仅关注局部输出而非全局代理行为。为此,我们提出一种面向过程的可解释性框架,揭示代理决策在时间、工具调用及架构边界上的演进过程。该框架围绕三个支柱构建:意图种子注入、推理监控与因果图界面,提供意图层级的运行状态数据,补充而非替代现有审查实践。我们在两类用户中验证其有效性:进行‘氛围编程’的普通用户获得原本被功能成功掩盖的结构可见性;专业开发者则在不增加负担的前提下获得更丰富的上下文支持用于代码审查。通过将认知漂移视为与代码质量同等重要的核心关切,本框架确保了代理监督仍具实质性意义。

原文摘要 · Abstract (English)

As autonomous coding agents become deeply embedded in software development workflows, their high operational velocity introduces a critical oversight challenge: the accumulating divergence between agentic actions and architectural intent. We term this process agentic entropy: a systemic drift that traditional code diff-based and HCXAI methods fail to capture, as they address local outputs rather than global agentic behaviour. To close this gap, we propose a process-oriented explainability framework that exposes how agentic decisions unfold across time, tool calls, and architectural boundaries. Built around three pillars (conformity seeding, reasoning monitoring, and a causal graph interface) our approach provides intent-level telemetry that complements, rather than replaces, existing review practices. We demonstrate its relevance across two user profiles: lay users engaged in vibe coding, who gain structural visibility otherwise masked by functional success; and professional developers, who gain richer contextual grounding for code review without increased overhead. By treating cognitive drift as a first-class concern alongside code quality, our framework supports the minimum level of human comprehension required for agentic oversight to remain substantive.

智能编程可解释性代理熵代码审查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。