arXiv:2502.17420cs.LGcs.AI2025-02ICML被引 100

发现大模型拒绝行为有多个独立机制,而非单一方向。

The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

论文配图:The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
图 1 · 摘自论文原文
  • 用梯度方法识别出多个独立的拒绝方向
  • 揭示拒绝由多维概念锥结构控制,非单一方向决定
  • 提出表征独立性新概念,适合安全对齐研究者

大型语言模型(LLMs)的安全对齐可能被对抗性输入绕过,但其攻击机制仍不明确。以往研究认为模型激活空间中存在单一拒绝方向决定是否拒绝请求。本文提出一种基于梯度的表示工程方法,发现存在多个独立的拒绝方向,甚至多维概念锥结构调控拒绝行为。此外,我们证明正交性并不等于干预下的独立性,提出了兼顾线性和非线性效应的表征独立性概念。通过该框架,我们识别出机制上独立的拒绝方向,证实了拒绝行为由多个不同机制驱动。该梯度方法可揭示这些机制,并为未来理解大模型提供基础。

原文摘要 · Abstract (English)

The safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood. Prior work suggests that a single refusal direction in the model's activation space determines whether an LLM refuses a request. In this study, we propose a novel gradient-based approach to representation engineering and use it to identify refusal directions. Contrary to prior work, we uncover multiple independent directions and even multi-dimensional concept cones that mediate refusal. Moreover, we show that orthogonality alone does not imply independence under intervention, motivating the notion of representational independence that accounts for both linear and non-linear effects. Using this framework, we identify mechanistically independent refusal directions. We show that refusal mechanisms in LLMs are governed by complex spatial structures and identify functionally independent directions, confirming that multiple distinct mechanisms drive refusal behavior. Our gradient-based approach uncovers these mechanisms and can further serve as a foundation for future work on understanding LLMs.

大模型安全拒绝机制表征独立性概念锥

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。