提出评估大模型最坏情况鲁棒性的新方法,发现多数防御几乎无效。
Towards the Worst-case Robustness of Large Language Models
- 用更强的白盒攻击上界分析最坏鲁棒性
- 证明多数确定性防御最坏情况鲁棒性接近0%
- 首次给出随机防御的理论下界,可验证任意攻击
近期研究揭示了大语言模型在对抗攻击下的脆弱性,攻击者可通过特定输入序列诱导有害、暴力、隐私或错误输出。本文研究其最坏情况鲁棒性,即是否存在导致不良输出的对抗样本。通过更强的白盒攻击上界估计,表明当前大多数确定性防御的最坏情况鲁棒性接近0%。提出一种通用紧致下界方法,基于分数背包或0-1背包求解器,用于界定所有随机防御的最坏情况鲁棒性。基于此,为若干已有经验防御提供理论下界。例如,对使用均匀核的平滑防御,可证明其在任意攻击下具有平均ℓ₀扰动2.02或平均后缀长度6.41的鲁棒性。
原文摘要 · Abstract (English)
Recent studies have revealed the vulnerability of large language models to adversarial attacks, where adversaries craft specific input sequences to induce harmful, violent, private, or incorrect outputs. In this work, we study their worst-case robustness, i.e., whether an adversarial example exists that leads to such undesirable outputs. We upper bound the worst-case robustness using stronger white-box attacks, indicating that most current deterministic defenses achieve nearly 0\% worst-case robustness. We propose a general tight lower bound for randomized smoothing using fractional knapsack solvers or 0-1 knapsack solvers, and using them to bound the worst-case robustness of all stochastic defenses. Based on these solvers, we provide theoretical lower bounds for several previous empirical defenses. For example, we certify the robustness of a specific case, smoothing using a uniform kernel, against \textit{any possible attack} with an average $\ell_0$ perturbation of 2.02 or an average suffix length of 6.41.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。