研究提示:攻击位置比攻击内容更重要,可能暴露安全评估漏洞。
Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models
- 发现攻击成功与否与对抗令牌位置密切相关,非仅靠后缀
- 调整提示中对抗词位置可显著提升攻击成功率
- 适合关注大模型安全评估的研究者和防御设计者
大型语言模型(LLMs)在多个领域广泛应用,亟需可靠的对齐安全机制。然而,通过对抗性提示绕过对齐的越狱攻击仍具挑战性。本文聚焦主流的贪婪坐标梯度(GCG)攻击,揭示了以往被忽视的攻击维度——对抗性标记在提示中的位置。以GCG为例,我们发现优化生成前缀而非后缀,以及在评估时改变对抗标记的位置,均会显著影响攻击成功率。结果表明当前安全评估存在关键盲区,必须考虑对抗标记的位置因素,才能全面评估大模型的鲁棒性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms. However, robustness remains challenging due to jailbreak attacks that bypass alignment via adversarial prompts. In this work, we focus on the prevalent Greedy Coordinate Gradient (GCG) attack and identify a previously underexplored attack axis in jailbreak attacks typically framed as suffix-based: the placement of adversarial tokens within the prompt. Using GCG as a case study, we show that both optimizing attacks to generate prefixes instead of suffixes and varying adversarial token position during evaluation substantially influence attack success rates. Our findings highlight a critical blind spot in current safety evaluations and underline the need to account for the position of adversarial tokens in the adversarial robustness evaluation of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。