用密码学机制确保AI无法突破安全底线,防极端威胁。
Governable AI: Provable Safety Under Extreme Threat Models
- 以密码学保障外部规则强制执行,不依赖内部约束
- 原型验证在高风险场景中实现不可绕过、不可篡改的安全防护
- 适合对安全性要求极高的关键系统,如自主决策平台
随着人工智能快速发展,其带来的安全风险日益严峻,尤其在可能引发系统性灾难的关键场景中。若AI失控、被操纵或主动规避安全机制,将导致严重后果。现有安全方法——如模型增强、价值对齐与人工干预——在面对具备极端动机和无限智能的AI时,存在根本性局限,无法保证安全。为此,本文提出可治理人工智能(GAI)框架,将安全机制从内部约束转向基于密码学的外部强制结构合规,该机制在既定威胁模型和公认密码假设下,计算上无法被破解,即便面对未来超智能AI亦然。GAI由规则执行模块(REM)、治理规则及可治理安全超级平台(GSSP)构成,实现全链路防护。REM执行治理底线规则,GSSP确保不可绕过、防篡改、防伪造,消除所有已知攻击路径。本文提供形式化安全证明,并通过原型在代表性高风险场景中验证其有效性。
原文摘要 · Abstract (English)
As AI rapidly advances, the security risks posed by AI are becoming increasingly severe, especially in critical scenarios, including those posing existential risks. If AI becomes uncontrollable, manipulated, or actively evades safety mechanisms, it could trigger systemic disasters. Existing AI safety approaches-such as model enhancement, value alignment, and human intervention-suffer from fundamental, in-principle limitations when facing AI with extreme motivations and unlimited intelligence, and cannot guarantee security. To address this challenge, we propose a Governable AI (GAI) framework that shifts from traditional internal constraints to externally enforced structural compliance based on cryptographic mechanisms that are computationally infeasible to break, even for future AI, under the defined threat model and well-established cryptographic assumptions.The GAI framework is composed of a simple yet reliable, fully deterministic, powerful, flexible, and general-purpose rule enforcement module (REM); governance rules; and a governable secure super-platform (GSSP) that offers end-to-end protection against compromise or subversion by AI. The decoupling of the governance rules and the technical platform further enables a feasible and generalizable technical pathway for the safety governance of AI. REM enforces the bottom line defined by governance rules, while GSSP ensures non-bypassability, tamper-resistance, and unforgeability to eliminate all identified attack vectors. This paper also presents a rigorous formal proof of the security properties of this mechanism and demonstrates its effectiveness through a prototype implementation evaluated in representative high-stakes scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。