arXiv:2507.20964cs.AIcs.CC2025-07AAAI被引 7

提出可证明可纠正的智能体框架,确保多步环境下的安全服从性。

Core Safety Values for Provably Corrigible Agents

  • 五类独立效用头分层加权,实现严格优先级控制
  • 即使学习误差ε,仍能保证安全属性不被违反且人类获益
  • 适用于对抗篡改场景,支持可验证的安全认证

我们提出了首个在部分可观测环境下对关机游戏具有完整形式化解决方案的可纠正性框架,并提供多步、部分可观测环境中的可证明保证。该框架包含五个结构分离的效用头——顺从性、开关访问保留、诚实性、基于信念扩展的可达成效用保持以实现低影响行为,以及有界任务奖励——通过严格的权重差距进行字典序组合。定理1证明了在部分可观测关机游戏中单轮可纠正性的精确保证;定理3将该保证扩展至多步自复制智能体,表明即使每个效用头的学习均方误差为ε,规划器也仅ε-次优,违反任一安全属性的概率仍受控,同时确保净人类利益。与宪法AI或RLHF/RLAIF等将所有规范合并为单一学习标量的方法不同,本方案的分离设计使得服从性与影响限制在激励冲突时仍能被证明占优。针对敌手可篡改智能体的情形,我们证明判断任意后攻击状态智能体是否会违反可纠正性是不可判定的(归约至停机问题),进而划出有限时域的‘可判定岛’,在此范围内安全性可在随机多项式时间内认证,并通过隐私保护、常数轮零知识证明验证。

原文摘要 · Abstract (English)

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *structurally separate* utility heads -- deference, switch-access preservation, truthfulness, low-impact behavior via a belief-based extension of Attainable Utility Preservation, and bounded task reward -- combined lexicographically by strict weight gaps. Theorem 1 proves exact single-round corrigibility in the partially observable off-switch game; Theorem 3 extends the guarantee to multi-step, self-spawning agents, showing that even if each head is *learned* to mean-squared error $\varepsilon$ and the planner is $\varepsilon$-sub-optimal, the probability of violating *any* safety property is bounded while still ensuring net human benefit. In contrast to Constitutional AI or RLHF/RLAIF, which merge all norms into one learned scalar, our separation makes obedience and impact-limits provably dominate even when incentives conflict. For settings where adversaries can modify the agent, we prove that deciding whether an arbitrary post-hack agent will ever violate corrigibility is undecidable by reduction to the halting problem, then carve out a finite-horizon "decidable island" where safety can be certified in randomized polynomial time and verified with privacy-preserving, constant-round zero-knowledge proofs.

可纠正性形式化安全强化学习智能体对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。