arXiv:2602.01587cs.CLcs.AI2026-02

通过噪声增强对齐实现大模型越狱攻击的可证明防御

Provable Defense Framework for LLM Jailbreaks via Noise-Augumented Alignment

  • 将安全防护从单次推理转为集成模型的统计稳定性
  • 梯度攻击成功率从84.2%降至1.2%,正常任务保留94.1%性能
  • 适用于需要高可信安全性的关键场景,如医疗与金融

大型语言模型仍易受自适应越狱攻击,现有基于经验的防御(如GCG)难以应对。本文提出一种可证明鲁棒性框架,将安全保证从单次推理转移至集成模型的统计稳定性。引入分层随机消融的认证语义平滑(CSS),通过将输入划分为不可变结构提示与可变载荷,利用超几何分布推导出严格的ℓ₀范数保证。为缓解稀疏上下文下的性能下降,采用噪声增强对齐调优(NAAT),使基础模型转化为语义去噪器。在Llama-3上的实验表明,该方法将基于梯度的攻击成功率从84.2%降至1.2%,同时保持94.1%的良性任务性能,显著优于字符级基线(性能降为74.3%)。该框架提供确定性安全证书,确保模型对所有可证明半径内的对抗变体保持鲁棒。

原文摘要 · Abstract (English)

Large Language Models (LLMs) remain vulnerable to adaptive jailbreaks that easily bypass empirical defenses like GCG. We propose a framework for certifiable robustness that shifts safety guarantees from single-pass inference to the statistical stability of an ensemble. We introduce Certified Semantic Smoothing (CSS) via Stratified Randomized Ablation, a technique that partitions inputs into immutable structural prompts and mutable payloads to derive rigorous lo norm guarantees using the Hypergeometric distribution. To resolve performance degradation on sparse contexts, we employ Noise-Augmented Alignment Tuning (NAAT), which transforms the base model into a semantic denoiser. Extensive experiments on Llama-3 show that our method reduces the Attack Success Rate of gradient-based attacks from 84.2% to 1.2% while maintaining 94.1% benign utility, significantly outperforming character-level baselines which degrade utility to 74.3%. This framework provides a deterministic certificate of safety, ensuring that a model remains robust against all adversarial variants within a provable radius.

大模型安全越狱防御可证明鲁棒噪声增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。