arXiv:2606.28639cs.LOcs.AI2026-06

AGI安全无法被验证,无论系统是否自我演化。

The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension

  • 用计算理论证明:无法完全、可靠、高效验证AGI行为安全。
  • 静态与动态安全验证均受不可解性限制,如图灵停机问题。
  • 任何监督者若能审计AGI,自身也需是通用AI,形成无限回归。

本文从数学上揭示了通用人工智能(AGI)安全性的双重极限:一是对固定系统的安全验证,二是对自修改后持续安全性的验证。在静态情形下,无论输入域无界(受Rice定理和哥德尔定理阻碍)或硬件配置有限(受Trakhtenbrot定理阻碍,表现为PSPACE难和co-RE完备性),都无法实现完全、可靠且高效的认证,导致‘可靠性-完备性-可计算性’三难困境,这是结构性而非统计性的必然。在动态情形下,将自修改形式化为可计算转移算子后,证明无法从当前已认证的安全状态判断下一阶段是否仍安全,这本质上是Rice定理的高阶应用,使静态与动态障碍成为同一根源的两面。这意味着持续认证仅可能存在于语义停止演化的系统中——即非通用系统。监督也无法规避此困境:能审计通用AGI的监督者本身必为通用智能,监督回归永无终止。三种实践风险(有限测试覆盖、受限推理时间、观察范围受限)本质相同:任何不拒绝正确证据的有界方案,都可能在每一步认证中维持虚假的安全表象,而实际属性始终被违反。这些结果赋予人工智能不可验证性以严格的数学内涵,表明其并非当前技术局限导致的工程难题,而是由与停机问题同源的计算规律所决定的表达力不变性。

原文摘要 · Abstract (English)

We establish the mathematical limits of AGI safety in two forms: verifying a fixed system, and verifying that a certified safety property persists once the system self-modifies. In the static case, no algorithm can certify a highly expressive AGI's safe behaviour infallibly, completely and tractably, whether over unbounded input domains (blocked by Rice's and Godel's theorems) or over all finite hardware configurations (blocked by Trakhtenbrot's theorem, which splits into a PSPACE-hardness barrier and a co-RE-completeness barrier), forcing a Soundness-Completeness-Tractability Trilemma as a structural, not statistical, necessity. In the dynamic case, we formalise self-modification as a computable transition operator and prove that no algorithm can determine, from a system's current certified safety, whether safety survives its next self-modification step: a result that reduces to Rice's Theorem one level up, making the static and dynamic barriers two faces of one obstruction. This forces an exclusive dichotomy: persistent certification is attainable only for systems that have stopped evolving semantically, i.e. only for narrow, not general, systems. Nor can the obstruction be delegated: any supervisor adequate to audit a general AGI is itself a general AGI, so the supervisory regress never terminates. Three practical risks (finite test coverage, bounded deliberation time, restricted observation) are one phenomenon: every bounded scheme that does not reject correct evidence admits an evolution trace it certifies at every stage while the property is persistently violated. These results give formal content to the unverifiability of AI, showing it is not an engineering target deferred by current limits but a structural tension, an Expressivity Invariant governed by the same computational laws as the Halting Problem and Rice's Theorem.

AGI安全计算理论不可验证性自修改

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。