预测激活调控的副作用,提前评估模型行为风险。
Forecasting Side Effects of Activation Steering

- 构建67种行为的交叉影响矩阵,分析副作用规律。
- 副作用普遍且不对称,目标行为决定影响程度。
- 基于原始模型表征可高精度预测副作用方向,适合安全审计。
激活调控通过向语言模型的隐藏层添加学习方向,实现无需重训练的行为调整。然而,该方法常引发其他行为的意外变化,难以安全部署。本文提出:能否在调控前预测这些副作用?我们通过对三个开源语言模型的67种行为构建交叉效应矩阵,发现副作用普遍存在、结构清晰且常呈非对称性,现有基于相似性的启发式方法无法解释。尽管复杂,副作用在调控前仍高度可预测:其幅度主要由目标行为决定,方向则可通过模型未调控前的表示进行预测,准确率显著高于基线。结果表明,激活调控具有系统性且可预测的副作用,支持主动安全审计与更明智的调控部署。
原文摘要 · Abstract (English)
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model's unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。