考虑采样过程的攻击能更真实评估大模型安全风险
Sampling-aware Adversarial Attacks Against Large Language Models
- 将攻击视为优化与采样资源分配问题,提升攻击效率
- 采样融入攻击后成功率最高提升37%,效率提高100倍
- 发现多数优化策略对有害输出影响有限,适合安全研究者
为保障大语言模型在大规模部署中的安全与鲁棒性,准确评估其对抗鲁棒性至关重要。现有攻击方法多针对单次贪婪生成中的有害响应,忽视了大模型固有的随机性,导致鲁棒性被高估。我们发现,在诱发有害响应的目标下,攻击过程中重复采样模型输出可有效补充提示优化,成为强大而高效的攻击手段。通过将攻击建模为优化与采样间的资源分配问题,我们实证确定了计算最优权衡,表明将采样融入现有攻击可使成功率最高提升37%,效率提升达两个数量级。我们进一步分析了攻击过程中输出有害性分布的变化,发现许多常见优化策略对输出有害性影响甚微。最后,我们提出一种基于熵最大化的无标签概念验证目标,展示了采样感知视角如何催生新型优化目标。总体而言,我们的研究揭示了采样在攻击中的关键作用,为规模化评估和增强大模型安全性提供了新路径。
原文摘要 · Abstract (English)
To guarantee safe and robust deployment of large language models (LLMs) at scale, it is critical to accurately assess their adversarial robustness. Existing adversarial attacks typically target harmful responses in single-point greedy generations, overlooking the inherently stochastic nature of LLMs and overestimating robustness. We show that for the goal of eliciting harmful responses, repeated sampling of model outputs during the attack complements prompt optimization and serves as a strong and efficient attack vector. By casting attacks as a resource allocation problem between optimization and sampling, we empirically determine compute-optimal trade-offs and show that integrating sampling into existing attacks boosts success rates by up to 37\% and improves efficiency by up to two orders of magnitude. We further analyze how distributions of output harmfulness evolve during an adversarial attack, discovering that many common optimization strategies have little effect on output harmfulness. Finally, we introduce a label-free proof-of-concept objective based on entropy maximization, demonstrating how our sampling-aware perspective enables new optimization targets. Overall, our findings establish the importance of sampling in attacks to accurately assess and strengthen LLM safety at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。