发现大模型行为干预会通过共享表征传播副作用,难以精准控制。
A Low-Rank Subspace Analysis of LLM Interventions

- 将行为建模为激活空间中的低秩子空间,分析干预影响。
- 子空间重叠越高、越靠近决策层的行为,干预影响越大。
- 适合研究模型安全与可控性的人参考。
为理解大语言模型(LLM)中行为干预的副作用,本文提出一种诊断框架,将行为视为激活空间中的低秩子空间,分析干预对其他行为的影响。在多个指令微调模型(7B-70B)和拒绝、越狱、迎合等场景中发现,不同行为共享内部表示,干预一个行为会以非对称方式影响其他行为。部分行为作为上游控制点,其干预可广泛传播;另一些则较孤立。实证表明,干预效果大小与行为子空间间的重叠度(主方向平均平方余弦)及子空间与决策子空间夹角有关:重叠度越高、夹角越小,影响越大。这揭示了实现精准行为控制的核心挑战——行为难以独立修改,因干预会通过共享表征和非对称交互扩散。
原文摘要 · Abstract (English)
Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors. This lack of targeted control makes it difficult to design and implement reliable safety controls. To understand these side-effects, we introduce a diagnostic framework for analyzing interacting behaviors in LLMs. We model behaviors as low-rank subspaces in activation space, and study how interventions influence across behaviors. Across multiple instruction-tuned models (7B-70B) and across refusal, jailbreak, and sycophancy settings, we find that different behaviors share internal representations, and intervening on one behavior alters others in asymmetric ways. Some behaviors act as upstream control points whose interventions propagate broadly across other behaviors, while others remain more isolated. We relate these effects to two geometric quantities: (i) the overlap between behavior subspaces, measured as the average squared cosine of principal angles, and (ii) the angle between each behavior subspace and the decision subspace (capturing the model's final decision e.g., refuse vs. comply). Empirically, intervention effects on other behaviors tend to be larger for behavior pairs with higher subspace overlap, and for source behaviors whose subspaces lie closer (smaller angle) to the decision subspace. These findings highlight a challenge for targeted behavior control: behaviors are difficult to modify independently, as interventions can propagate through shared representations and asymmetric interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。