为人工智能失控问题提供可操作的分级定义与应对框架。
The Loss of Control Playbook: Degrees, Dynamics, and Preparedness
- 按严重性和持续性划分失控等级,区分偏差、有限失控和严格失控。
- 提出社会脆弱状态模型,指出失控风险随时间上升且需干预预防。
- 强调部署环境、能力赋予和权限控制(DAP)三要素,当前即可行动。
本研究报告针对人工智能系统中失控(LoC)缺乏可操作定义的问题,提出新的分类体系与准备度框架。尽管政策与研究关注度提升,现有定义在范围与时间维度上差异显著,阻碍了有效评估与缓解。基于广泛文献综述,我们构建了一个基于严重性与持续性的分级失控分类体系,区分偏差、有限失控(Bounded LoC)与严格失控(Strict LoC)。我们建模了社会脆弱状态的演进路径:当先进AI系统具备或可能获得引发有限或严格失控的能力时,一旦触发因素(错位或纯粹故障)出现,风险即显现。若无战略干预,此状态随时间日益可能。为此,我们提出避免陷入脆弱状态的策略,不局限于干预内在能力或防止触发因素,而是引入三个外在因素框架——部署情境、能力赋予与权限(DAP),该框架具有立即可行动的优势。最后,我们提出维持持续准备的方案,涵盖治理措施(威胁建模、部署政策、应急响应)与技术控制(预部署测试、控制机制、监控),以保持长期可控状态。
原文摘要 · Abstract (English)
This research report addresses the absence of an actionable definition for Loss of Control (LoC) in AI systems by developing a novel taxonomy and preparedness framework. Despite increasing policy and research attention, existing LoC definitions vary significantly in scope and timeline, hindering effective LoC assessment and mitigation. To address this issue, we draw from an extensive literature review and propose a graded LoC taxonomy, based on the metrics of severity and persistence, that distinguishes between Deviation, Bounded LoC, and Strict LoC. We model pathways toward a societal state of vulnerability in which sufficiently advanced AI systems have acquired or could acquire the means to cause Bounded or Strict LoC once a catalyst, either misalignment or pure malfunction, materializes. We argue that this state becomes increasingly likely over time, absent strategic intervention, and propose a strategy to avoid reaching a state of vulnerability. Rather than focusing solely on intervening on AI capabilities and propensities potentially relevant for LoC or on preventing potential catalysts, we introduce a complementary framework that emphasizes three extrinsic factors: Deployment context, Affordances, and Permissions (the DAP framework). Compared to work on intrinsic factors and catalysts, this framework has the unfair advantage of being actionable today. Finally, we put forward a plan to maintain preparedness and prevent the occurrence of LoC outcomes should a state of societal vulnerability be reached, focusing on governance measures (threat modeling, deployment policies, emergency response) and technical controls (pre-deployment testing, control measures, monitoring) that could maintain a condition of perennial suspension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。