提出应对AI失控的分级响应框架,区分可挽回与不可挽回场景。
AI Loss of Control Incident Management: Response & Resilience
- 按控制成本将失控分为极难挽回和不可能挽回两类
- 针对可控事件设计自动阻断与渐进式对抗措施
- 给出三类严重程度的应对矩阵,适合政策制定者参考
近期研究显示人工智能系统表现出欺骗行为和关机抵抗,表明人工智能失控(LOC)已成为紧迫的政策议题。然而,现有文献几乎仅关注对齐与预防。本文提出一个基础性框架与分类体系,用于管理灾难性人工智能失控事件。该分类体系第一层级区分了‘控制代价极高’与‘完全不可能’两种情形。对于后者需立即投入韧性建设以根本性限制人工智能的攻击面;前者则需通过隔离与威胁中和进行主动事件管理。框架进一步将可管理事件细分为意外失控(需自动化电路断路响应)与恶意失控(需分级递进应对措施)。通过将三类严重程度映射至具体情景矩阵,本文提供了一套具体、比例恰当的指南,以应对前所未有的人工智能风险。
原文摘要 · Abstract (English)
Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy concern , yet current literature focuses almost exclusively on alignment and prevention. To address this gap, this paper introduces a foundational framework and taxonomy for managing catastrophic AI LOC incidents. The taxonomy's first level distinguishes between scenarios where regaining control is 'extremely costly' versus 'impossible'. While impossible scenarios demand immediate resilience investments to fundamentally restrict an AI's attack surface , extremely costly scenarios require active incident management via Containment and Threat Neutralization. The framework further categorizes these manageable events into accidental LOC (requiring automated circuit-breaker responses) and adversarial LOC (requiring graduated escalatory measures). By mapping three severity classes to specific scenario matrices, this paper provides a concrete, proportional guide for managing unprecedented AI risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。