通过校准建模安全对齐的扭曲,提升大模型越狱成功率。
Jailbreaking LLMs via Calibration
- 将越狱视为预测聚合问题,提出梯度偏移优化策略。
- 在gpt-oss-120b上实现更高攻击成功率与更低越狱代价。
- 适用于对抗测试、安全评估等前沿模型研究场景。
大型语言模型的安全对齐常导致其输出分布与原始预对齐数据分布之间产生系统性偏差。本文将该偏差建模为预对齐分布的系统性扭曲,将弱到强越狱问题转化为预测聚合任务,并推导出基于损失诱导对偶空间中梯度偏移的最优聚合策略。我们证明了对数算术越狱方法是交叉熵损失下的特例,并推导出适用于其他合理损失函数的更广泛聚合规则。此外,提出一种新型混合聚合规则。在多个红队测试基准和数学实用性任务上,使用前沿模型进行评估表明,该方法在攻击成功率方面优于现有技术,尤其在经过强化安全防护的gpt-oss-120b模型上表现更优,且显著降低‘越狱代价’。
原文摘要 · Abstract (English)
Safety alignment in Large Language Models (LLMs) often creates a systematic discrepancy between a model's aligned output and the underlying pre-aligned data distribution. We propose a framework in which the effect of safety alignment on next-token prediction is modeled as a systematic distortion of a pre-alignment distribution. We cast Weak-to-Strong Jailbreaking as a forecast aggregation problem and derive an optimal aggregation strategy characterized by a Gradient Shift in the loss-induced dual space. We show that logit-arithmetic jailbreaking methods are a special case of this framework under cross-entropy loss, and derive a broader family of aggregation rules corresponding to other proper losses. We also propose a new hybrid aggregation rule. Evaluations across red-teaming benchmarks and math utility tasks using frontier models demonstrate that our approach achieves superior Attack Success Rates and lower "Jailbreak Tax" compared with existing methods, especially on the safety-hardened gpt-oss-120b.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。