激活引导在对抗性文本攻击下严重失效,影响模型行为控制的可靠性。
Adversarial Robustness of Activation Steering in Large Language Models
- 通过注入预计算向量实现无训练行为控制
- 对抗攻击使引导效果下降64%,置信度降至0.25以下
- 现有层选择方法对扰动敏感,难以实际应用
激活引导是一种无需训练的控制大语言模型行为的方法,通过在推理时向残差流注入预计算的方向向量实现。然而其在真实输入变化下的鲁棒性尚未被研究。本文首次系统评估了激活引导在对抗性文本扰动下的表现,覆盖四种提取方法、三种攻击策略、六个来自Anthropic Model-Written Evaluation Dataset的人格设定,以及参数量从1.5B到30B的五种模型。攻击在所有设置中均成功:方向鲁棒性下降最多达64%,攻击后置信度普遍降至0.25以下,几乎所有可引导输入的引导强度均下降。层选择同样脆弱,自动识别的最优层在扰动下最多偏移17层,加剧了向量层面的失效。从对抗样本中提取向量可部分恢复PCA和MD方法在中大型模型上的可引导性,但无法准确找到优化后的最佳层,限制了缓解措施的实际价值。结果表明,激活引导的脆弱性是结构性的,而非方法特异性,当前层选择策略不足以支持真实部署。
原文摘要 · Abstract (English)
Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's residual stream at inference time. Yet its robustness to realistic input variation remains unstudied. We present the first systematic evaluation of activation steering robustness under adversarial text perturbations on the inputs, covering four extraction methods, three attack strategies, six personas from Anthropic Model-Written Evaluation Dataset, and five models ranging from 1.5B to 30B parameters. Attacks succeed broadly across all settings: directional robustness drops by up to 64%, post-attack confidence collapses near or below 0.25 across all methods and models, and steering strength degrades on nearly every steerable input. Layer selection is equally fragile, with the optimal layer identified by an automated method on clean inputs shifting by up to 17 positions under perturbation, a failure that compounds the vector-level breakdown. Extracting vectors from adversarially perturbed inputs partially recovers steerability for PCA and MD on mid-to-large models, but they consistently fail to locate the improved optimal layer, limiting the practical benefit of this mitigation. Together, these findings reveal that the brittleness of activation steering is structural rather than method-specific, and that current layer selection strategies are not robust enough for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。