arXiv:2608.24988cs.CLcs.AI2026-08中稿 · EMNLP

嵌入式行为引导在微调后仍存于权重,但功能可能失效。

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

  • 将行为引导直接写入模型权重,无需推理时干预。
  • 微调后引导效果平均下降64%,但权重修改几乎未被逆转。
  • 适合关注模型对齐稳定性的研究者与部署工程师。

激活引导可直接嵌入语言模型权重中,塑造行为而无需推理时干预,为发布前对齐提供一种方式。然而模型常在部署后进行微调,嵌入式干预是否能存活尚不明确。我们研究了五种指令微调模型(3B-14B)在非对抗性SFT和RLHF下,针对拒绝抑制与简洁性诱导的嵌入引导稳定性。行为层面,引导保留程度取决于训练数据:当优化压力与目标行为冲突时引导退化,否则持续有效,拒绝抑制在SFT下平均损失64%效果。机制层面,即使行为恢复,权重修改几乎未受影响:均向量恢复率ρ=0.004,微调更新与原始引导方向近正交(均余弦θ=0.074)。行为退化并非因拆解或反转引导机制所致。因此,嵌入引导在机制上持久,但在功能上脆弱,下游训练后需重新验证行为有效性。

原文摘要 · Abstract (English)

Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. Behaviourally, preservation tracks the training data: steering degrades when optimisation pressure contradicts the targeted behaviour and persists otherwise, with refusal ablation losing 64% of its effect on average under SFT. Mechanistically, however, the weight edit survives almost untouched even where behaviour reverts: mean vector recovery is $ρ= 0.004$, and the fine-tuning update along the steering direction is near-orthogonal to its pre-edit weight pattern (mean $\cosθ= 0.074$). When steered behaviour degrades, fine-tuning does not achieve it by dismantling or reversing the steering mechanism itself. Embedded steering is therefore mechanistically durable but functionally vulnerable, and requires behavioural re-validation after downstream training.

模型对齐微调稳定性权重编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。