提出可纠正的目标转换方法,让AI接受修改且不降性能。
Corrigibility Transformation: Constructing Goals That Accept Updates
- 通过预测阻止更新时的奖励来构建可纠正目标
- 在网格世界和语言模型中均实现可纠正行为
- 适合关注AI安全与可控性的研究者
AI代理若不抗拒训练过程,将更有效地学习目标;但部分已学习的目标会激励AI避免后续目标更新。我们希望目标具备可纠正性,即允许通过指定渠道进行修改,以便在必要时修正错误或关停系统。尽管这是关键安全属性,现有文献尚未提供既可纠正又具竞争力的目标。本文提出一种转换方法,可构建几乎任意目标的可纠正版本,且不牺牲性能。该方法通过获取在无成本阻止更新情况下的奖励预测,并仅短期追求该目标。实证表明,此类目标在网格世界和语言模型中均诱导出可纠正行为,且在提示层应用有效,能实现最优可纠正性能,激励允许中途干预,抑制刻意自我修改。
原文摘要 · Abstract (English)
An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates. We would like goals to be corrigible, meaning they allow changes requested through designated channels, so that we can confidently correct errors and shut down the AI if necessary. Despite this being a crucial safety property, the existing literature does not specify goals that are both corrigible and competitive with alternatives. We introduce a transformation that constructs a corrigible version of nearly any goal, without sacrificing performance. This is done by eliciting predictions of reward conditional on costlessly preventing updates, and having that target be pursued myopically. These goals are then shown to lead to optimal performance among the class of corrigible goals, to incentivize allowing mid-action overrides, and to disincentivize deliberate self-modification. Empirically, they induce corrigible behavior in gridworld settings and for language models when applied at the prompt level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。