arXiv:2605.12963cs.AI2026-05

AI安全不能只靠外部控制,必须内建安全机制。

Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements

论文配图:Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements
图 1 · 摘自论文原文
  • 用控制理论证明:外部控制失效后,依赖它的安全策略必然失败
  • 若要持续安全,系统必须具备内在的安全目标与稳定性
  • 适合关注AI长期安全的学者和开发者阅读

随着人工智能系统能力提升,安全策略不仅要降低当前风险,还必须能在外部控制失效后仍维持安全。本文运用控制理论,在结构层面分析外部强制安全策略是否可行。首先,在可达性等前提下,证明一旦系统影响超出有限外部控制范围,任何依赖外部持续约束的安全策略都无法持续保障安全——这是整个外部策略类别的结构性失败。其次,若仍有可行策略存在,则所有剩余策略必须是内在的。文中提出四项结构要求:安全不能依赖外部持续控制;初始目标必须安全兼容;目标在自我修改中需保持稳定;能力增长时安全必须持续维持。论文未提出完整方案,而是通过形式化推导,明确哪些策略被排除,哪些条件必须满足,为普遍担忧提供理论框架。

原文摘要 · Abstract (English)

As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could sustain safety once external control can no longer reliably constrain system behavior. This paper addresses that problem by using control theory to clarify, at a structural level, whether externally enforced safety-sustaining strategies can succeed and, if not, what any alternative strategy would have to satisfy in order to be viable. It establishes two main results. First, under explicit premises including a reachability condition, it proves a class-wide external impossibility result: once the system's effects exceed what bounded external control can counteract, no strategy that depends in any degree on continued external enforcement can sustain AI safety. This failure is structural across the entire externally enforced class rather than contingent on any particular strategy. Second, it establishes a conditional class-level necessity result: if at least one candidate safety-sustaining strategy remains after that elimination, then all such remaining strategies must be intrinsic. It then states four structural requirements for viability: safety may not depend on continued external enforcement; the system's terminal objective must be safety-compatible when first formed; that objective must remain stable under self-modification; and safety must continue to be preserved as capability grows. The paper does not propose a complete strategy for sustaining AI safety. Its contribution is to give formal structure to a widely held concern about the limits of external control. It does so by deriving explicit conditional results that identify which safety-sustaining strategies are ruled out and what any remaining strategies must satisfy.

AI安全控制理论内在安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。