arXiv:2606.12442cs.CYcs.AI2026-06

厘清人工智能失控的本质:控制是设定与达成目标的能力。

Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It

  • 以目标设定与实现定义控制,构建可操作的控制框架。
  • 指出即使非超智能AI也可能导致人类不同程度的失控。
  • 适合关注AI治理、系统安全与风险防范的研究者与决策者。

当前,人工智能失控风险在公众讨论中备受关注,学术界、前沿实验室乃至政府均展开广泛探讨。然而现有文献对失控概念的基础极为薄弱,缺乏对‘控制’本身的明确定义。本文旨在填补这一空白:首先基于‘目标设定与达成’建立控制的实用定义,并融合控制论、管理控制与控制理论等领域的基础概念,分析控制主体、目标设定能力、控制回路、所需多样性及目标对齐等要素。在此基础上,探讨控制如何丧失,以及人工智能如何促成这种丧失,并提出维护控制的具体建议。研究发现,即便人工智能未达到超智能水平,人类个体与群体也可能因其行为而失去不同程度的控制——我们所定义的失控场景早已存在且持续发生。

原文摘要 · Abstract (English)

At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discourse present among academics, frontier labs, and even governments. However, in the existing literature, the concept seems to rest on surprisingly weak foundations, where even those that discuss loss of control extensively do not first establish what control is and what exactly is being lost. Our paper aims to address these gaps. We establish a working definition of control by anchoring it to the "setting and getting of goals". Then, we discuss various aspects of control, built on foundational concepts from related fields like cybernetics, management control, and control theory. This includes who (or what) can be in control, and the things they require to be in control, such as the ability to set goals, having a functional control loop, having requisite variety, and having sufficient goal alignment. Once a framework for control is established, we then discuss how control can be lost, how AIs can contribute to such loss of control, and offer relevant recommendations for how one can maintain control. One interesting consequence of our work is that humanity, as individuals and as groups, can lose varying degrees of control as a result of AI behaviour that is far below the level of superintelligence; the potential for loss of control scenarios (as we define them) already exist, and have existed for a long time.

AI治理控制论风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。