arXiv:2511.08379cs.AIcs.LG2025-11AAAI被引 9

用自组织映射发现多个拒绝方向,更好抑制大模型拒答行为

SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models

  • 用自组织映射从有害提示中提取多个拒绝方向,替代单一方向
  • 删去多个方向后拒答率下降37.6%,优于单向基线和越狱攻击
  • 适合研究模型安全机制或对抗性攻击的学者参考

拒答是指安全对齐语言模型拒绝有害或不道德提示的功能行为。近期研究将拒答编码为模型隐空间中的单一方向,例如通过有害与无害提示表示中心点的差值计算。然而,越来越多证据表明,大模型中的概念常以低维流形形式嵌入高维隐空间。受此启发,我们提出一种新方法,利用自组织映射(SOM)提取多个拒答方向。首先证明SOM可推广先前的均值差技术;随后在有害提示表示上训练SOM以识别多个神经元,通过减去无害表示中心点,得到一组表达拒答概念的多方向向量。在广泛实验中验证,移除多个方向比单方向基线及专用越狱算法更有效抑制拒答行为,效果显著。最后分析了该方法的机制含义。

原文摘要 · Abstract (English)

Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model's latent space; e.g., computed as the difference between the centroids of harmful and harmless prompt representations. However, emerging evidence suggests that concepts in LLMs often appear to be encoded as a low-dimensional manifold embedded in the high-dimensional latent space. Motivated by these findings, we propose a novel method leveraging Self-Organizing Maps (SOMs) to extract multiple refusal directions. To this end, we first prove that SOMs generalize the prior work's difference-in-means technique. We then train SOMs on harmful prompt representations to identify multiple neurons. By subtracting the centroid of harmless representations from each neuron, we derive a set of multiple directions expressing the refusal concept. We validate our method on an extensive experimental setup, demonstrating that ablating multiple directions from models' internals outperforms not only the single-direction baseline but also specialized jailbreak algorithms, leading to an effective suppression of refusal. Finally, we conclude by analyzing the mechanistic implications of our approach.

模型安全拒答机制自组织映射可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。