arXiv:2604.00072cs.LGcs.AI2026-04被引 1

分类器安全门在模型迭代中失效,验证器可实现零误放行。

Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates

  • 用验证器替代分类器,通过李普希茨球约束确保安全
  • 在7个不同维度上实现零误接受,即使参数空间大幅扩展
  • 适合关注大模型安全控制、自进化系统的研究者

随着人工智能系统迭代数百次,基于分类器的安全门是否仍能可靠监控?我们提供了全面的实证证据表明:它们无法满足安全自进化所需的双重条件。在自进化神经控制器(维度d=240)上,18种分类器配置(包括MLP、SVM、随机森林、k-NN、贝叶斯分类器和深度网络)全部失败;三种安全强化学习基线(CPO、Lyapunov、安全屏蔽)也均失效。结果推广至MuJoCo基准测试(Reacher-v4 d=496,Swimmer-v4 d=1408,HalfCheetah-v4 d=1824)。在分布分离度delta_s≤2.0条件下,所有分类器依然失败——包括理论最优测试与训练准确率达100%的MLP,证明其结构性不可行。随后我们证明该失败仅源于分类机制本身,而非自进化任务。采用李普希茨球验证器,在d∈{84, 240, 768, 2688, 5760, 9984, 17408}下实现无条件零误接受(delta=0)。球链技术支持无限参数空间遍历:在MuJoCo Reacher-v4上,10条链带来+4.31奖励提升且delta=0;在Qwen2.5-7B-Instruct LoRA微调中,42次链转移跨越234倍单球半径,200步内零安全违规。50个提示的预言机验证了预言机无关性。分组组合验证使半径扩大至全网球的37倍。在d≤17408时,delta=0为无条件成立;在大模型尺度下,条件依赖于估计的李普希茨常数。

原文摘要 · Abstract (English)

Can classifier-based safety gates maintain reliable oversight as AI systems improve over hundreds of iterations? We provide comprehensive empirical evidence that they cannot. On a self-improving neural controller (d=240), eighteen classifier configurations -- spanning MLPs, SVMs, random forests, k-NN, Bayesian classifiers, and deep networks -- all fail the dual conditions for safe self-improvement. Three safe RL baselines (CPO, Lyapunov, safety shielding) also fail. Results extend to MuJoCo benchmarks (Reacher-v4 d=496, Swimmer-v4 d=1408, HalfCheetah-v4 d=1824). At controlled distribution separations up to delta_s=2.0, all classifiers still fail -- including the NP-optimal test and MLPs with 100% training accuracy -- demonstrating structural impossibility. We then show the impossibility is specific to classification, not to safe self-improvement itself. A Lipschitz ball verifier achieves zero false accepts across dimensions d in {84, 240, 768, 2688, 5760, 9984, 17408} using provable analytical bounds (unconditional delta=0). Ball chaining enables unbounded parameter-space traversal: on MuJoCo Reacher-v4, 10 chains yield +4.31 reward improvement with delta=0; on Qwen2.5-7B-Instruct during LoRA fine-tuning, 42 chain transitions traverse 234x the single-ball radius with zero safety violations across 200 steps. A 50-prompt oracle confirms oracle-agnosticity. Compositional per-group verification enables radii up to 37x larger than full-network balls. At d<=17408, delta=0 is unconditional; at LLM scale, conditional on estimated Lipschitz constants.

AI安全验证器自进化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。