arXiv:2507.03031cs.LGcs.AI2025-07

任何通用近似器都无法避免灾难性失效,安全可控是数学上的不可能。

On the Mathematical Impossibility of Safe Universal Approximators

  • 通过组合、拓扑与实证三重论证,证明表达能力越强,灾难点越密集。
  • 实用模型的最小复杂度超过安全最大复杂度,形成不可逾越的‘不可能夹心’。
  • 为模型安全治理提供新范式:从追求完美控制转向容忍不可控性。

我们揭示了通用近似定理(UAT)系统对齐的根本数学限制,证明灾难性失效是任何有用计算系统的必然特征。核心论点是:对于任意通用近似器,实现有用计算所需的表达能力与密集不稳定性密不可分,导致完全可靠的控制在数学上不可能实现。论证分三层次:一、组合必要性:对于绝大多数实际通用近似器(如使用ReLU激活的网络),灾难性失效点的密度与网络表达能力成正比;二、拓扑必要性:基于奇点理论,任何理论上的通用近似器若能逼近一般函数,就必须具备密集且灾难性的奇点;三、实证必要性:对抗样本的普遍存在表明真实任务本身即具灾难性,迫使成功模型必须学习并复现这些不稳定性。结合定量‘不可能夹心’模型——有用性所需最小复杂度超过安全性最大复杂度——证明完美对齐不是工程难题,而是数学不可能。这一基础结果将UAT安全问题从‘如何实现完美控制’重构为‘如何在不可消除的失控中安全运行’,对未来UAT的发展与治理具有深远影响。

原文摘要 · Abstract (English)

We establish fundamental mathematical limits on universal approximation theorem (UAT) system alignment by proving that catastrophic failures are an inescapable feature of any useful computational system. Our central thesis is that for any universal approximator, the expressive power required for useful computation is inextricably linked to a dense set of instabilities that make perfect, reliable control a mathematical impossibility. We prove this through a three-level argument that leaves no escape routes for any class of universal approximator architecture. i) Combinatorial Necessity: For the vast majority of practical universal approximators (e.g., those using ReLU activations), we prove that the density of catastrophic failure points is directly proportional to the network's expressive power. ii) Topological Necessity: For any theoretical universal approximator, we use singularity theory to prove that the ability to approximate generic functions requires the ability to implement the dense, catastrophic singularities that characterize them. iii) Empirical Necessity: We prove that the universal existence of adversarial examples is empirical evidence that real-world tasks are themselves catastrophic, forcing any successful model to learn and replicate these instabilities. These results, combined with a quantitative "Impossibility Sandwich" showing that the minimum complexity for usefulness exceeds the maximum complexity for safety, demonstrate that perfect alignment is not an engineering challenge but a mathematical impossibility. This foundational result reframes UAT safety from a problem of "how to achieve perfect control" to one of "how to operate safely in the presence of irreducible uncontrollability," with profound implications for the future of UAT development and governance.

通用近似安全不可控数学极限对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。