模板区域让大模型安全机制变脆弱,易被攻击绕过
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
- 发现大模型安全机制依赖固定模板区域的语义信息
- 实验证明该机制普遍存在,导致模型易被推理时攻击破解
- 建议将安全机制与模板解耦,提升防御能力
大型语言模型(LLMs)的安全对齐仍存在漏洞,其初始行为极易被简单攻击绕过。由于现有模型在输入指令与初始输出之间填充固定模板是常见做法,我们提出假设:该模板区域是导致安全脆弱性的关键因素——模型的安全决策过度依赖模板区域的聚合信息,从而影响其安全行为。我们称此现象为模板锚定安全对齐。本文通过大量实验验证了该现象在多种对齐大模型中普遍存在。机制分析表明,这种依赖关系使模型在遭遇推理时的越狱攻击时尤为脆弱。此外,我们证明将安全机制从模板区域中解耦,可有效缓解越狱攻击风险。我们呼吁未来研究发展更鲁棒的安全对齐技术,降低对模板区域的依赖。
原文摘要 · Abstract (English)
The safety alignment of large language models (LLMs) remains vulnerable, as their initial behavior can be easily jailbroken by even relatively simple attacks. Since infilling a fixed template between the input instruction and initial model output is a common practice for existing LLMs, we hypothesize that this template is a key factor behind their vulnerabilities: LLMs' safety-related decision-making overly relies on the aggregated information from the template region, which largely influences these models' safety behavior. We refer to this issue as template-anchored safety alignment. In this paper, we conduct extensive experiments and verify that template-anchored safety alignment is widespread across various aligned LLMs. Our mechanistic analyses demonstrate how it leads to models' susceptibility when encountering inference-time jailbreak attacks. Furthermore, we show that detaching safety mechanisms from the template region is promising in mitigating vulnerabilities to jailbreak attacks. We encourage future research to develop more robust safety alignment techniques that reduce reliance on the template region.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。