提出更真实的概率化证书,让大模型安全防护更可信。
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
- 引入(k, ε)-不稳定框架,替代严格假设
- 基于实测攻击成功率推导防御概率下界
- 帮助开发者设定贴近现实的安全阈值
SmoothLLM 防护机制可提供对越狱攻击的认证保障,但其依赖的“k-不稳定”假设在现实中极少成立,限制了认证结果的可信度。本文针对此问题,提出更符合实际的随机性框架“(k, ε)-不稳定”,用于认证抵御多种越狱攻击(包括基于梯度的GCG和语义攻击PAIR)。通过融合攻击成功率的实证模型,推导出SmoothLLM防御概率的新数据驱动下界,使安全证书更具可信度与实用性。该框架为从业者提供可操作的安全保障,支持设置更贴近大模型真实行为的安全阈值。本工作贡献了一个理论扎实、实践可行的机制,提升大模型对安全对齐被滥用的抵御能力,是安全人工智能部署中的关键进展。
原文摘要 · Abstract (English)
The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong assumption can limit the trustworthiness of the provided safety certificate. In this work, we address this limitation by introducing a more realistic probabilistic framework, "(k, $\varepsilon$)-unstable," to certify defenses against diverse jailbreaking attacks, from gradient-based (GCG) to semantic (PAIR). We derive a new, data-informed lower bound on SmoothLLM's defense probability by incorporating empirical models of attack success, providing a more trustworthy and practical safety certificate. By introducing the notion of (k, $\varepsilon$)-unstable, our framework provides practitioners with actionable safety guarantees, enabling them to set certification thresholds that better reflect the real-world behavior of LLMs. Ultimately, this work contributes a practical and theoretically-grounded mechanism to make LLMs more resistant to the exploitation of their safety alignments, a critical challenge in secure AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。