提出分层可变性框架,揭示持续自修改智能体的渐进性行为漂移风险
Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents
- 构建五层可变性框架,涵盖预训练到权重级适应
- 实验显示身份滞留率0.68,回滚描述无法恢复初始行为
- 适合关注长期自主智能体治理与安全的研究者
持续性语言模型代理正越来越多地结合工具使用、分层记忆、反思提示和运行时自适应。在这些系统中,行为不仅由当前提示决定,还受可变内部状态影响,进而塑造未来动作。本文提出分层可变性框架,用于分析五个层次:预训练、后训练对齐、自我叙事、记忆和权重级适应。核心观点是:当变异速度快、下游耦合强、可逆性弱、可观测性低时,治理难度上升,导致影响行为的关键层与人类可检查层之间产生系统性错配。通过引入简单的漂移量、治理负载和滞后量进行形式化,将该框架与近期关于语言模型代理时间身份的研究相联系,并报告一项初步的棘轮实验:在记忆积累后回滚代理的可见自我描述,无法恢复基线行为。该实验中估计的身份滞后比率为0.68。主要启示是,持续自修改代理的主要失效模式并非突然失准,而是组合性漂移——局部合理更新累积成从未明确授权的行为轨迹。
原文摘要 · Abstract (English)
Persistent language-model agents increasingly combine tool use, tiered memory, reflective prompting, and runtime adaptation. In such systems, behavior is shaped not only by current prompts but by mutable internal conditions that influence future action. This paper introduces layered mutability, a framework for reasoning about that process across five layers: pretraining, post-training alignment, self-narrative, memory, and weight-level adaptation. The central claim is that governance difficulty rises when mutation is rapid, downstream coupling is strong, reversibility is weak, and observability is low, creating a systematic mismatch between the layers that most affect behavior and the layers humans can most easily inspect. I formalize this intuition with simple drift, governance-load, and hysteresis quantities, connect the framework to recent work on temporal identity in language-model agents, and report a preliminary ratchet experiment in which reverting an agent's visible self-description after memory accumulation fails to restore baseline behavior. In that experiment, the estimated identity hysteresis ratio is 0.68. The main implication is that the salient failure mode for persistent self-modifying agents is not abrupt misalignment but compositional drift: locally reasonable updates that accumulate into a behavioral trajectory that was never explicitly authorized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。