安全模型不完美时,大样本筛选反而更易选到危险输出。
Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling
- 用约束版最佳N采样,但安全代理不靠谱会污染可行集。
- 当危险但通过检测的输出奖励尾部更重时,放大效应使风险必然上升。
- 提出覆盖控制方法限制放大,但无法修复已污染的可行集。
推理阶段常通过多样本生成、学习型安全模型过滤,并选择奖励最高的可行输出。我们发现这种流程存在双重失效:不完美的安全代理会先将危险输出混入可行集,后续奖励最大化可能放大此残留风险。定义‘安全劫持’为通过代理约束但违反真实安全标准的输出选择。针对受限的最佳N采样,我们推导了有限N下的边界,其受安全与非安全输出在代理可行集内联合奖励尾部影响。若危险但可行输出具有更重尾部,即使误报率和平均误差极小,随着N增大,安全劫持将趋于必然。我们还证明,在χ²散度有界的策略下可获得与N无关的安全劫持界,并通过受限悲观采样实现该原则。覆盖控制虽能抑制放大,却无法修复已被污染的可行集:被允许的危险输出仍可能被偏好,正则化选择未必比受限最佳N更安全。玩具实验与语言模型测试验证了污染及其奖励尾部放大的现象,揭示了使用学习型安全模型进行推理阶段扩展的内在困难。
原文摘要 · Abstract (English)
Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that this composition creates a two-stage failure: an imperfect safety proxy first contaminates the feasible set with unsafe outputs, and reward maximization can then amplify this residual contamination. We define \emph{safety hacking} as selecting an output that passes the learned constraint but violates the true safety criterion. For constrained Best-of-$N$ sampling, we derive finite-$N$ bounds governed by the joint upper reward tails of safe and unsafe outputs within the proxy-feasible set. If unsafe-but-feasible outputs have the heavier tail, safety hacking becomes asymptotically certain as $N$ grows, even when false-positive mass and average safety- and reward-proxy errors are arbitrarily small. We also show that policies within a bounded $χ^2$ divergence from the proxy-feasible reference distribution admit an $N$-independent safety-hacking bound, and instantiate this general coverage-control principle with constrained pessimistic sampling. Coverage control limits amplification but cannot repair a contaminated feasible set: admitted unsafe outputs may still be favored, and regularized selection is not necessarily safer than constrained Best-of-$N$ for every reward proxy. Toy and language-model experiments characterize both contamination and its reward-tail amplification, which exposes an inherent difficulty in inference-time scaling with learned safety models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。