提出抗干扰且能自适应的AI安全框架,应对动态环境中的未知风险。
\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World
- 用生物免疫机制启发,让安全成为持续对抗学习过程。
- 结合快慢模型与安全风洞仿真,实现对已知和未知威胁的双重防护。
- 适合关注AI长期安全与通用智能演进的研究者与开发者。
本文针对快速发展的AI能力与滞后的安全进展之间的鸿沟提出新思路。现有方法分为‘让AI变安全’(事后对齐,脆弱且被动)和‘让安全的AI’(强调内在安全性,但难以应对开放环境中的未知风险)。为此,我们提出‘安全共进化’作为新范式,借鉴生物免疫系统,使安全成为动态、对抗性且持续的学习过程。为实现这一愿景,提出 exttt{R$^2$AI}——抗干扰且具韧性的AI框架,融合快慢安全模型、通过安全风洞进行对抗模拟与验证,并建立持续反馈回路,推动安全与能力协同演化。该框架为动态环境中持续保障安全提供了可扩展、主动的路径,兼顾近期漏洞与长远存在性风险,助力AI向通用智能(AGI)和超智能(ASI)演进。
原文摘要 · Abstract (English)
In this position paper, we address the persistent gap between rapidly growing AI capabilities and lagging safety progress. Existing paradigms divide into ``Make AI Safe'', which applies post-hoc alignment and guardrails but remains brittle and reactive, and ``Make Safe AI'', which emphasizes intrinsic safety but struggles to address unforeseen risks in open-ended environments. We therefore propose \textit{safe-by-coevolution} as a new formulation of the ``Make Safe AI'' paradigm, inspired by biological immunity, in which safety becomes a dynamic, adversarial, and ongoing learning process. To operationalize this vision, we introduce \texttt{R$^2$AI} -- \textit{Resistant and Resilient AI} -- as a practical framework that unites resistance against known threats with resilience to unforeseen risks. \texttt{R$^2$AI} integrates \textit{fast and slow safe models}, adversarial simulation and verification through a \textit{safety wind tunnel}, and continual feedback loops that guide safety and capability to coevolve. We argue that this framework offers a scalable and proactive path to maintain continual safety in dynamic environments, addressing both near-term vulnerabilities and long-term existential risks as AI advances toward AGI and ASI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。