arXiv:2506.10304cs.AIcs.CC2025-06被引 3

AI安全与能力的矛盾源于根本逻辑困境,无法通过技术突破解决。

The Alignment Trap: Complexity Barriers

  • 从五个数学层面证明安全AI存在不可逾越的障碍
  • 安全策略在策略空间中占比为零,验证其安全性属计算难题
  • 适合关注AI对齐本质挑战的研究者和政策制定者

本文认为,人工智能对齐不仅困难,更根植于一个基本的逻辑矛盾。我们首先提出枚举悖论:使用机器学习正是因为我们无法枚举所有安全规则,但让机器学习变安全却需要来自这一不可能枚举的示例。该悖论通过五项独立的数学证明(“不可能性支柱”)得到验证。主要结果包括:(1) 几何不可能性:安全策略集合测度为零,是将无限维世界上下文要求投影到有限维模型的必然结果;(2) 计算不可能性:即使允许非零误差容限,验证策略安全性也是coNP完全问题;(3) 统计不可能性:训练所需的安全数据(大量罕见灾难实例)在逻辑上自相矛盾,无法获得;(4) 信息论不可能性:安全规则包含的信息量超过任何可行网络可存储的极限;(5) 动态不可能性:提升能力的优化过程与安全目标梯度通常反向。这些结果共同表明,追求高能力且安全的AI并非克服技术障碍的问题,而是面对一系列相互关联的根本壁垒。论文最后提出一个战略三难困境,迫使该领域作出抉择。核心定理的Lean4形式化验证正在进展中。

原文摘要 · Abstract (English)

This paper argues that AI alignment is not merely difficult, but is founded on a fundamental logical contradiction. We first establish The Enumeration Paradox: we use machine learning precisely because we cannot enumerate all necessary safety rules, yet making ML safe requires examples that can only be generated from the very enumeration we admit is impossible. This paradox is then confirmed by a set of five independent mathematical proofs, or "pillars of impossibility." Our main results show that: (1) Geometric Impossibility: The set of safe policies has measure zero, a necessary consequence of projecting infinite-dimensional world-context requirements onto finite-dimensional models. (2) Computational Impossibility: Verifying a policy's safety is coNP-complete, even for non-zero error tolerances. (3) Statistical Impossibility: The training data required for safety (abundant examples of rare disasters) is a logical contradiction and thus unobtainable. (4) Information-Theoretic Impossibility: Safety rules contain more incompressible, arbitrary information than any feasible network can store. (5) Dynamic Impossibility: The optimization process for increasing AI capability is actively hostile to safety, as the gradients for the two objectives are generally anti-aligned. Together, these results demonstrate that the pursuit of safe, highly capable AI is not a matter of overcoming technical hurdles, but of confronting fundamental, interlocking barriers. The paper concludes by presenting a strategic trilemma that these impossibilities force upon the field. A formal verification of the core theorems in Lean4 is currently in progress.

AI对齐逻辑悖论安全屏障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。