提出多轮安全认证框架,有效抵御逐步攻击
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

- 用状态对抗马尔可夫决策过程建模对话安全
- 多轮安全概率衰减从指数级降至更慢速率
- 适用于大模型安全验证,尤其抗渐进式越狱
大型语言模型易受多轮越狱攻击,攻击者通过逐步操纵对话上下文实现突破。现有认证鲁棒性方法仅适用于单轮输入;直接组合会导致安全边界随轮次呈指数级恶化。本文提出多轮认证鲁棒性(MTCR)框架,将对话安全建模为状态对抗马尔可夫决策过程,并定义k轮认证鲁棒性为在k次对抗轮次下最坏情况的安全概率。MTCR包含:(i) 通过嵌入空间模式分解实现组合认证,获得比简单乘法更紧的下界;(ii) (α,β)-安全持续性机制,将衰减速率从$\underline{p}^{k}$提升至$β^k$(其中$β> \underline{p}$),并提供可解释的预测时域;(iii) 与信息论上界匹配,证明其紧致性;(iv) 一个统一算法整合上述成果。六种大模型在ε有界攻击和Crescendo风格攻击下的实验表明,实际安全性能始终高于认证边界。
原文摘要 · Abstract (English)
Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines $k$-turn certified robustness as the worst-case safety probability across $k$ adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) $(α,β)$-safety persistence, improving the degradation rate from $\underline{p}^{k}$ to $β^k$ (with $β> \underline{p}$) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under $ε$-bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。