arXiv:2607.20581cs.CRcs.AI2026-07

研究扰动提示的几何结构,揭示大模型安全漏洞的内在机制。

Geometric Configurations of Perturbed Jailbreak Prompts

论文配图:Geometric Configurations of Perturbed Jailbreak Prompts
图 1 · 摘自论文原文
  • 分析扰动提示在模型表征空间中的分布形态
  • 发现仅少数词元与合规响应显著相关
  • 为防御提示劫持提供可解释性依据

扰动技术不断演化,将失败的越狱提示转为成功攻击,严重威胁大模型安全。本文研究 Qwen-2.5-1.5B/-3B/-7B-Instruct 和 Llama-3.2-1B/-3B/-3.1-8B-Instruct 系列小参数模型中字符串级扰动输入的内部表征。选取两个表征空间:最后一层最后一个标记的嵌入空间,以及前50个预测词元的概率空间。前者按拼写和格式区分提示,后者虽看似复杂,实则近似一维。在拒绝主导的回答集合中,两个空间均未发现行为超平面。仅在1.5B Qwen模型中,下一个词元"Sure"与合规回答显著相关;在1$ Llama模型中,","和"ĊĊ"也表现出显著关联。

原文摘要 · Abstract (English)

Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "ĊĊ" in the 1$ Llama model, display a significant association with a compliant-labeled answer.

提示劫持模型安全表征分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。