分析智能体追求权力的理论基础,质疑其必然性
Instrumental convergence and power-seeking

- 论证权力追求依赖于强版本工具收敛假说
- 指出现有辩护均未能充分支撑该假说
- 适合关注AI长期风险与治理的研究者阅读
近年来,人们对人工智能可能对人类生存构成重大威胁日益担忧。其中一个主要依据是,人工智能代理可能具有追求权力的倾向,从而获取控制权并削弱人类的自主性。本文指出,这一担忧建立在强版本的工具收敛假说之上。我探讨了该假说的主要辩护,并认为它们均未充分证明其足够强的形式以支撑权力追求论点。文章进一步讨论了该结论对长期主义、人工智能治理以及研究智能体风险方法论的影响。
原文摘要 · Abstract (English)
Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is that artificial agents may be power-seeking, aiming to acquire power and in the process disempowering humanity. I show how the argument from power-seeking rests on a strong version of a claim known as the instrumental convergence thesis. I explore leading defenses of the instrumental convergence thesis and argue that none establishes the thesis in a strong enough form to ground the argument from power-seeking. I discuss implications for longtermism, the governance of artificial intelligence, and the methodology of studying risks posed by artificial agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。