用多样语法结构攻击大模型安全对齐,成功率超96%。
StructTransform: A Scalable Attack Surface for Safety-Aligned Large Language Models
- 通过自然语言意图的多种语法格式编码实施攻击
- 简单攻击达90%成功率,组合策略突破96%且零拒绝
- 揭示现有防御依赖词级模式,适合安全研究者参考
本文提出一系列针对大语言模型对齐的安全结构转换攻击,通过将自然语言意图编码为从简单结构到全新由大模型生成的语法空间等多种形式。大规模评估显示,最简单的攻击在严格对齐模型(如Claude 3.5 Sonnet)上已实现近90%的成功率。通过结合结构变换与已有内容变换的自适应策略,进一步将攻击成功率提升至96%以上,且无任何拒绝。我们探索了多种结构格式,包括完全由大模型生成的新语法,发现其易生成且具高成功率,表明防御难度显著。最后,我们构建基准并测试现有安全对齐防御,结果显示多数方法失败率达100%。结果表明,当前对齐机制多依赖词级模式而忽视有害概念,亟需更深入研究。作为案例,我们演示了攻击者如何利用该方法轻松生成可绕过检测的恶意软件样本和欺诈短信语料库。
原文摘要 · Abstract (English)
In this work, we present a series of structure transformation attacks on LLM alignment, where we encode natural language intent using diverse syntax spaces, ranging from simple structure formats and basic query languages (e.g., SQL) to new novel spaces and syntaxes created entirely by LLMs. Our extensive evaluation shows that our simplest attacks can achieve close to a 90% success rate, even on strict LLMs (such as Claude 3.5 Sonnet) using SOTA alignment mechanisms. We improve the attack performance further by using an adaptive scheme that combines structure transformations along with existing content transformations, resulting in over 96% ASR with 0% refusals. To generalize our attacks, we explore numerous structure formats, including syntaxes purely generated by LLMs. Our results indicate that such novel syntaxes are easy to generate and result in a high ASR, suggesting that defending against our attacks is not a straightforward process. Finally, we develop a benchmark and evaluate existing safety-alignment defenses against it, showing that most of them fail with 100% ASR. Our results show that existing safety alignment mostly relies on token-level patterns without recognizing harmful concepts, highlighting and motivating the need for serious research efforts in this direction. As a case study, we demonstrate how attackers can use our attack to easily generate a sample malware and a corpus of fraudulent SMS messages, which perform well in bypassing detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。