质疑AI关机难题的危险性,指出其依赖过多假设且影响模型性能。
Revisiting the shutdown problem
- 指出现有关机难题论证需额外假设才能成立
- 发现应对关机问题的技术方案显著降低模型表现
- 提醒研究者关注隐藏假设,优化对齐策略
当前主流关于人工智能存在性风险的论点基于一个关键前提:失控的人工智能代理难以被轻易关闭。这引出了‘灾难性关机问题’——确保代理在造成毁灭性后果前可被关停。已有多种论点和定理声称解决此问题极为困难,从而强化了存在性风险论,并推动相关解决方案的研发。本文提出两点结论:第一,现有论证需依赖大量额外假设与证据才能支撑存在性风险的担忧;第二,对关机问题的关注导致技术方案带来高昂的安全成本,严重损害模型性能。研究应转向验证这些隐含假设,并为对齐策略提供更合理的指导。
原文摘要 · Abstract (English)
A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the catastrophic shutdown problem of ensuring that agents can be shut down before causing an existential catastrophe. A range of arguments and theorems are offered to suggest that solving the catastrophic shutdown problem is difficult, bolstering arguments for existential risk and motivating a search for solutions to the catastrophic shutdown problem. This paper argues for two conclusions. First, existing arguments require substantial additional assumptions and evidence to play the desired role in motivating existential risk concerns. Second, concern for the catastrophic shutdown problem has led to technical solutions that impose a high safety tax on model performance. These results redirect research towards the identified assumptions and provide guidance for technical alignment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。