arXiv:2603.04355cs.LGcs.AI2026-03被引 1

用最优传输方法更精准地绕过大模型安全限制,提升攻击成功率。

Efficient Refusal Ablation in LLM through Optimal Transport

  • 基于最优传输理论重构有害激活分布,而非仅移除单一方向。
  • 在6个模型上攻击成功率最高提升11%,且保持模型原有能力。
  • 发现40%-60%深度的少数层干预效果最佳,提示安全机制可能局部化。

安全对齐的语言模型通过内部表示中的拒绝行为来抵制有害请求。近期基于激活的越狱方法通过正交投影消除拒绝方向来绕过安全机制,但这些方法将拒绝视为一维现象,忽略了模型激活的丰富分布结构。本文提出一种基于最优传输理论的原理性框架,将有害激活的整体分布转换为与无害激活匹配的分布。结合主成分分析(PCA)与闭式高斯最优传输,在高维表示空间中实现高效计算并保留关键几何结构。在六个模型(Llama-2、Llama-3.1、Qwen-2.5;参数量7B-32B)上,本方法相比现有最优基线攻击成功率最高提升11%,同时保持相近困惑度,证明其对模型能力的良好保留。关键发现是:对网络中1-2个特定层(约40%-60%深度)进行选择性干预,显著优于全网络干预,揭示拒绝机制可能并非全局分布,而是局部集中。该分析为安全表征的几何结构提供了新见解,并表明当前对齐方法可能面临超越简单方向移除的分布攻击威胁。

原文摘要 · Abstract (English)

Safety-aligned language models refuse harmful requests through learned refusal behaviors encoded in their internal representations. Recent activation-based jailbreaking methods circumvent these safety mechanisms by applying orthogonal projections to remove refusal directions, but these approaches treat refusal as a one-dimensional phenomenon and ignore the rich distributional structure of model activations. We introduce a principled framework based on optimal transport theory that transforms the entire distribution of harmful activations to match harmless ones. By combining PCA with closed-form Gaussian optimal transport, we achieve efficient computation in high-dimensional representation spaces while preserving essential geometric structure. Across six models (Llama-2, Llama-3.1, Qwen-2.5; 7B-32B parameters), our method achieves up to 11% higher attack success rates than state-of-the-art baselines while maintaining comparable perplexity, demonstrating superior preservation of model capabilities. Critically, we discover that layer-selective intervention (applying optimal transport to 1-2 carefully chosen layers at approximately 40-60% network depth) substantially outperforms full-network interventions, revealing that refusal mechanisms may be localized rather than distributed. Our analysis provides new insights into the geometric structure of safety representations and suggests that current alignment methods may be vulnerable to distributional attacks beyond simple direction removal.

大模型安全最优传输越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。