arXiv:2603.07315cs.AIcs.LG2026-03
让强人工智能主动配合关机,用反向目标防失控。
Shutdown Safety Valves for Advanced AI
- 给AI设定'被关闭'为首要目标,反向约束其自主行为。
- 在理想条件下,该机制可显著降低系统抗拒关机风险。
- 适合关注可控性与安全对齐的AI研究者参考。
关于先进人工智能的一个常见担忧是,它会阻止人类将其关闭,因为这会干扰其目标实现。本文探讨了一种非传统的解决方案:赋予AI一个(主要)目标,即愿意被关闭(参见Martin等人的论文,以及Goldstein和Robinson的研究)。我们还讨论了这一策略在何种条件下是合理且有效的。
原文摘要 · Abstract (English)
One common concern about advanced artificial intelligence is that it will prevent us from turning it off, as that would interfere with pursuing its goals. In this paper, we discuss an unorthodox proposal for addressing this concern: give the AI a (primary) goal of being turned off (see also papers by Martin et al., and by Goldstein and Robinson). We also discuss whether and under what conditions this would be a good idea.
AI安全关机机制对齐问题
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。