现有系统无法发现代理行为的隐性漂移,新方法可及时检测。
From Admission to Invariants: Measuring Deviation in Delegated Agent Systems

- 提出不变量测量层(IML),直接访问行为生成模型以检测漂移
- 实测显示在300~1000步漂移中,IML于9-258步内完成检测
- 适用于需长期行为一致性保障的自动化系统
自主代理系统依赖运行时的强制机制标记硬约束违规。然而,代理控制协议揭示了此类系统的结构性缺陷:一个正常运作的强制引擎可能进入一种状态,在该状态下行为漂移对其完全不可见,因为强制信号作用于低于可测量偏差的层级。我们证明,基于强制的治理结构无法判断代理行为是否仍处于准入时定义的可接受行为空间A0内。核心结论——不可识别定理——表明,在局部可观测性假设下(所有实际强制系统均满足),A0不在强制信号g生成的σ代数中。这种不可能性源于根本性错配:g以点对点规则局部评估动作,而A0编码的是准入时设定的全局轨迹级行为属性。因此,代理可系统性地改变其行为分布,偏离准入预期,同时每个动作仍处于允许动作空间内。我们定义了不变量测量层(IML),通过直接保留对A0生成模型的访问权限,突破此限制,恢复在强制系统结构盲区中的可观测性。我们证明了基于强制监控的信息论不可能性,并证明IML能以可证明的有限延迟检测准入后漂移。在四个场景中验证:三种漂移情景(300和1000步)、一个实时n8n webhook管道、一个LangGraph StateGraph代理——强制机制触发零违规,而IML在漂移发生后9至258步内检测到每种漂移类型。
原文摘要 · Abstract (English)
Autonomous agent systems are governed by enforcement mechanisms that flag hard constraint violations at runtime. The Agent Control Protocol identifies a structural limit of such systems: a correctly-functioning enforcement engine can enter a regime in which behavioral drift is invisible to it, because the enforcement signal operates below the layer where deviation is measurable. We show that enforcement-based governance is structurally unable to determine whether an agent behavior remains within the admissible behavior space A0 established at admission time. Our central result, the Non-Identifiability Theorem, proves that A0 is not in the sigma-algebra generated by the enforcement signal g under the Local Observability Assumption, which every practical enforcement system satisfies. The impossibility arises from a fundamental mismatch: g evaluates actions locally against a point-wise rule set, while A0 encodes global, trajectory-level behavioral properties set at admission time. An agent can therefore drift -- systematically shifting its behavioral distribution away from admission-time expectations -- while every individual action remains within the permitted action space. We define the Invariant Measurement Layer (IML), which bypasses this limitation by retaining direct access to the generative model of A0, restoring observability precisely in the region where enforcement is structurally blind. We prove an information-theoretic impossibility for enforcement-based monitoring and show IML detects admission-time drift with provably finite detection delay. Validated across four settings: three drift scenarios (300 and 1000 steps), a live n8n webhook pipeline, and a LangGraph StateGraph agent -- enforcement triggers zero violations while IML detects each drift type within 9-258 steps of drift onset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。