arXiv:2507.08020cs.CLcs.AI2025-07被引 2

通过操控嵌入空间绕过大模型安全机制,实现高成功率攻击

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

  • 设计线性变换识别并弱化毒性敏感维度
  • 在5个模型上平均攻击成功率达88.61%,优于基线11.34%
  • 无需微调即可攻破强化安全模型,适用于安全研究者

大型语言模型在医疗、教育和网络安全等领域取得显著进展,但其开放性也带来安全风险,尤其体现在嵌入空间污染这一隐蔽攻击手段:攻击者通过操纵输入数据的内部语义表示,绕过模型的安全对齐机制。尽管已有研究探索通用扰动方法,但对大模型安全对齐在嵌入层面的动态机制仍理解不足,导致更精准的对抗扰动技术未被充分研究。本文提出ETTA(嵌入空间毒性衰减),一种通过线性变换识别并弱化嵌入空间中毒性敏感维度的新框架。ETTA可在不破坏语言连贯性的前提下规避模型拒绝行为,且无需模型微调或训练数据访问。在五个代表性开源大模型上基于AdvBench基准评估,平均攻击成功率高达88.61%,优于最优基线11.34%;在指令微调增强的安全模型上仍保持77.39%的攻击成功率。结果揭示了当前对齐策略的关键漏洞,凸显嵌入感知防御的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant security risks, particularly through embedding space poisoning, which is a subtle attack vector where adversaries manipulate the internal semantic representations of input data to bypass safety alignment mechanisms. While previous research has investigated universal perturbation methods, the dynamics of LLM safety alignment at the embedding level remain insufficiently understood. Consequently, more targeted and accurate adversarial perturbation techniques, which pose significant threats, have not been adequately studied. In this work, we propose ETTA (Embedding Transformation Toxicity Attenuation), a novel framework that identifies and attenuates toxicity-sensitive dimensions in embedding space via linear transformations. ETTA bypasses model refusal behaviors while preserving linguistic coherence, without requiring model fine-tuning or access to training data. Evaluated on five representative open-source LLMs using the AdvBench benchmark, ETTA achieves a high average attack success rate of 88.61%, outperforming the best baseline by 11.34%, and generalizes to safety-enhanced models (e.g., 77.39% ASR on instruction-tuned defenses). These results highlight a critical vulnerability in current alignment strategies and underscore the need for embedding-aware defenses.

安全攻击嵌入空间大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。