揭示大模型控制向量的本质不可识别性,挑战行为干预的可解释性假设。
On the Non-Identifiability of Steering Vectors in Large Language Models
- 通过正交扰动验证控制向量在行为上等效,证明其不可唯一确定。
- 不同模型和属性下,正交扰动效果相近,输出差异极小。
- 该现象为几何固有特性,适用于多种输入分布,影响对齐可靠性。
激活操控方法广泛用于控制大语言模型行为,常被解读为揭示了有意义的内部表征。这一解读依赖于操控方向可识别且唯一可恢复的假设。我们发现,在白盒单层访问条件下,操控向量因存在大量行为不可区分的干预方式而根本上不可识别。实证表明,正交扰动在多个模型和特质中达到近似等效的操控效果,输出层面差异可忽略。通过奇异值分解计算激活协方差矩阵的零空间维度,验证了等效性在操作相关范围内的稳健性。关键的是,这种不可识别性是鲁棒的几何性质,不随提示分布变化而消失。这些发现揭示了可解释性的根本局限,强调仅靠行为测试无法实现可靠的对齐干预,需引入结构约束。
原文摘要 · Abstract (English)
Activation steering methods are widely used to control large language model (LLM) behavior and are often interpreted as revealing meaningful internal representations. This interpretation assumes that steering directions are identifiable and uniquely recoverable from input-output behavior. We show that, under white-box single-layer access, steering vectors are fundamentally non-identifiable due to large equivalence classes of behaviorally indistinguishable interventions. Empirically, we find that orthogonal perturbations achieve near-equivalent efficacy with negligible effect sizes across multiple models and traits, with pre-trained semantic classifiers confirming equivalence at the output level. We estimate null-space dimensionality via SVD of activation covariance matrices and validate that equivalence holds robustly throughout the operationally relevant steering range. Critically, we show that non-identifiability is a robust geometric property that persists across diverse prompt distributions. These findings reveal fundamental interpretability limits and highlight the need for structural constraints beyond behavioral testing to enable reliable alignment interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。