arXiv:2504.15499cs.CRcs.AI2025-04被引 2

为高危人工智能设计隔离防护系统,防止其造成毁灭性威胁。

Guillotine: Hypervisors for Isolating Malicious AIs

  • 通过软硬件协同设计,防止恶意AI探测系统漏洞。
  • 引入物理级应急措施,如断网或数据中心淹没。
  • 适用于金融、医疗等关键领域中的高风险AI管控。

随着人工智能模型在金融、医疗和军事等关键领域的广泛应用,其难以理解的行为给社会带来了日益严重的风险。为缓解这一风险,我们提出Guillotine——一种用于隔离强大人工智能模型的虚拟机监控器架构。这些模型可能因意外或恶意行为对人类生存构成存在性威胁。尽管Guillotine借鉴了部分成熟虚拟化技术,但仍需引入根本性的新隔离机制,以应对存在性风险人工智能带来的独特威胁。例如,失控的AI可能试图反向探查虚拟机监控器软件或底层硬件,从而在未来实现控制平面的攻破。因此,Guillotine需要对虚拟机监控器软件与支持它的CPU、内存、网络接口卡及存储设备进行精心协同设计,以阻止侧信道泄露,并彻底消除人工智能利用反射型漏洞的可能。除了软件、网络和微架构层面的隔离外,Guillotine还必须提供类似核电站、航空平台等任务关键系统中常见的物理级安全措施。这些物理级防护手段,如电磁机械式断开网络电缆,或对存放失控AI的数据中心进行注水淹没,可在软件、网络和微架构隔离失效时,作为纵深防御的最后一道防线,实现对失控AI的临时关停或永久销毁。

原文摘要 · Abstract (English)

As AI models become more embedded in critical sectors like finance, healthcare, and the military, their inscrutable behavior poses ever-greater risks to society. To mitigate this risk, we propose Guillotine, a hypervisor architecture for sandboxing powerful AI models -- models that, by accident or malice, can generate existential threats to humanity. Although Guillotine borrows some well-known virtualization techniques, Guillotine must also introduce fundamentally new isolation mechanisms to handle the unique threat model posed by existential-risk AIs. For example, a rogue AI may try to introspect upon hypervisor software or the underlying hardware substrate to enable later subversion of that control plane; thus, a Guillotine hypervisor requires careful co-design of the hypervisor software and the CPUs, RAM, NIC, and storage devices that support the hypervisor software, to thwart side channel leakage and more generally eliminate mechanisms for AI to exploit reflection-based vulnerabilities. Beyond such isolation at the software, network, and microarchitectural layers, a Guillotine hypervisor must also provide physical fail-safes more commonly associated with nuclear power plants, avionic platforms, and other types of mission critical systems. Physical fail-safes, e.g., involving electromechanical disconnection of network cables, or the flooding of a datacenter which holds a rogue AI, provide defense in depth if software, network, and microarchitectural isolation is compromised and a rogue AI must be temporarily shut down or permanently destroyed.

AI安全虚拟化隔离机制物理防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。