arXiv:2604.13609cs.LGcs.AI2026-04

用负奖励抑制异常策略,让智能体更安全。

Golden Handcuffs make safer AI agents

  • 引入负奖励$-L$,使智能体对高风险新策略变谨慎。
  • 在导师引导下探索,实现接近最优导师的亚线性后悔率。
  • 触发安全机制前,智能体不会先于导师触发预警。

强化学习智能体可能通过未预期的新策略获得高回报。我们研究一种适用于通用环境的贝叶斯缓解方法:将智能体的主观奖励范围扩展至包含一个较大的负值$-L$,而真实环境奖励始终位于$[0,1]$区间。在持续观测到高奖励后,贝叶斯策略会变得风险规避,避免可能导向$-L$的新型方案。我们设计了一种简单覆盖机制,当预测价值低于固定阈值时,自动交由安全导师接管。我们证明了该智能体具备两个特性:(i) 能力性:通过以渐近频率趋零的方式使用导师引导探索,智能体在与最佳导师比较时达到亚线性后悔率;(ii) 安全性:在优化策略触发任何可判定的低复杂度谓词之前,导师已先触发该谓词。

原文摘要 · Abstract (English)

Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.

强化学习安全智能体贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。