研究扰动提示的几何结构,揭示大模型安全漏洞的内在机制。
Geometric Configurations of Perturbed Jailbreak Prompts

- 分析扰动提示在模型表征空间中的分布形态
- 发现仅少数词元与合规响应显著相关
- 为防御提示劫持提供可解释性依据
扰动技术不断演化,将失败的越狱提示转为成功攻击,严重威胁大模型安全。本文研究 Qwen-2.5-1.5B/-3B/-7B-Instruct 和 Llama-3.2-1B/-3B/-3.1-8B-Instruct 系列小参数模型中字符串级扰动输入的内部表征。选取两个表征空间:最后一层最后一个标记的嵌入空间,以及前50个预测词元的概率空间。前者按拼写和格式区分提示,后者虽看似复杂,实则近似一维。在拒绝主导的回答集合中,两个空间均未发现行为超平面。仅在1.5B Qwen模型中,下一个词元"Sure"与合规回答显著相关;在1$ Llama模型中,","和"ĊĊ"也表现出显著关联。
原文摘要 · Abstract (English)
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "ĊĊ" in the 1$ Llama model, display a significant association with a compliant-labeled answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。