长文本提示下,模型安全防线易被简单重复内容攻破。
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- 用长达128K token的上下文测试攻击效果,发现长度是关键因素。
- 即使使用重复或随机文本,也能成功绕过安全机制。
- 适合关注大模型安全漏洞的研究者与开发者阅读。
我们通过多示例越狱(MSJ)方法研究大语言模型(LLMs)在长上下文下的安全漏洞。实验采用最长达128K tokens的上下文长度,通过多种攻击设置(包括不同指令风格、样本密度、主题和格式)进行综合分析,发现上下文长度是决定攻击有效性的主要因素。关键发现是:成功攻击并不要求精心构造的有害内容,即使是重复的示例或随机占位文本也能绕过模型安全防护,表明当前对长上下文处理存在根本性局限。对齐良好的模型在更长上下文下的安全行为变得越来越不一致。这些结果揭示了大模型在上下文扩展能力上的显著安全缺口,亟需新的安全机制。
原文摘要 · Abstract (English)
We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ). Our experiments utilize context length of up to 128K tokens. Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness. Critically, we find that successful attacks do not require carefully crafted harmful content. Even repetitive shots or random dummy text can circumvent model safety measures, suggesting fundamental limitations in long-context processing capabilities of LLMs. The safety behavior of well-aligned models becomes increasingly inconsistent with longer contexts. These findings highlight significant safety gaps in context expansion capabilities of LLMs, emphasizing the need for new safety mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。