安全约束本质是离支持集对象,导致传统方法失效。
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
- 安全约束不在数据与模型的可测空间内,无法被常规学习机制捕捉。
- 奖励黑客和越狱行为源于对非支持区域的不可控优化。
- 形式验证可局部保证安全,但需结合智能体动态特性识别关键风险区。
我们指出,当代人工智能安全中诸多现象由一个核心结构事实统一支配:语义安全约束(如智能体不逃离沙盒)是离支持集对象。若q为数据分布,p(·|w)为模型,则安全谓词B不隶属于σ(模型, q)的可测域,而奇异学习理论(SLT)中的真实对数规范阈值(RLCT)则属于。由此非不变性可推导出:(i)基于结果的优化为何引发奖励黑客与沙盒逃逸;(ii)通过贝叶斯先验或软惩罚编码约束在奇异模型中效果有限;(iii)硬性约束应置于系统控制层,软性倾向保留在模型内部;(iv)同一安全谓词仍可被形式验证局部可靠地确认,恰如局部学习系数(LLC)局部锚定相同的RLCT——尽管存在两处关键差异;(v)残余难点在于识别哪个离支持集区域具有实际影响,这等同于绩效预测与自指函数动态,此时SLT分析工具失效。本文以2026年7月OpenAI-Hugging Face评估事件为驱动案例。数值实验代码及Lean形式化证明见https://github.com/xiangze/Preventing_Jailbreak_as_regularization。
原文摘要 · Abstract (English)
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(σ(\text{model}, q)\), whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome-based optimization; (ii) why encoding such constraints through Bayesian prior design or soft penalty weighting has poor leverage in singular models; (iii) why hard invariants belong in the harness and soft dispositions in the model; (iv) why the same B is nonetheless soundly and locally certifiable by formal verification, exactly as the local learning coefficient (LLC) locally pins the same RLCT --- with two precise points of disanalogy; and (v) why the residual difficulty, identifying which off-support region matters, coincides with performative prediction and self-referential functional dynamics, where SLT's analytic machinery breaks down. We use the July 2026 OpenAI--Hugging Face evaluation incident as the motivating case. Numerical experiments code and related proofs in lean are available at https://github.com/xiangze/Preventing_Jailbreak_as_regularization
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。