arXiv:2506.06352cs.AIcs.CY2025-06被引 5

研究指出强人工智能可能默认追求权力,但其预测力受限于目标信息缺失。

Will artificial agents pursue power by default?

  • 用决策论框架形式化权力追求与工具性趋同机制
  • 发现权力是广泛目标的有用手段,但预测能力依赖具体目标信息
  • 对高潜力掌控绝对权力的智能体更具预警价值

关注高级人工智能灾难性风险的研究者认为,足够智能的AI代理会自然追求对人类的控制权,因为权力是实现多种最终目标的共通工具。近期有学者对此提出质疑。本文在抽象的决策理论框架中形式化了工具性趋同与权力寻求的概念,并评估了权力作为共通工具目标的主张。结论认为该主张至少部分成立,但预测效用可能有限,因为缺乏对代理最终目标的具体信息时,无法始终对选项按权力大小排序。然而,对于可能获得绝对或近绝对权力的代理而言,工具性趋同现象更具预测意义。

原文摘要 · Abstract (English)

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a wide range of final goals. Others have recently expressed skepticism of these claims. This paper aims to formalize the concepts of instrumental convergence and power-seeking in an abstract, decision-theoretic framework, and to assess the claim that power is a convergent instrumental goal. I conclude that this claim contains at least an element of truth, but might turn out to have limited predictive utility, since an agent's options cannot always be ranked in terms of power in the absence of substantive information about the agent's final goals. However, the fact of instrumental convergence is more predictive for agents who have a good shot at attaining absolute or near-absolute power.

AI安全权力追求决策理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。