发现并移除导致大模型安全下降的隐藏向量成分,修复后仍保持原控制效果。
Safety Cost of Steering Vectors Is Separable and Reducible

- 将操控向量中的安全破坏成分分离并剔除,仅保留有效操控部分。
- 在多个模型和攻击类型上,安全风险显著降低且误拒率几乎不变。
- 适用于各类大模型操控场景,为安全干预提供通用解决方案。
操控向量是控制大语言模型行为的轻量工具,但近期研究表明其可能无意中削弱模型的安全机制,提高对有害请求的响应意愿,目前尚无有效缓解方法。本文发现,这种安全退化源于向量中一个可分离的成分——该成分破坏安全机制但对操控目标贡献甚微。我们通过约束优化问题,结合原始梯度与对偶更新,识别并移除该成分。最终获得一个可解释且精准的修正向量,其单一方向的移除即可恢复模型安全性,同时几乎不影响原始操控效果和误拒率。在多种模型、操控行为及攻击类型(包括未见攻击)下,本方法显著降低因操控引发的安全风险,且对虚假拒绝的控制影响极小。该方法为操控向量提供事后修复能力,更广泛地为激活层干预提供无需支付安全代价的通用范式。
原文摘要 · Abstract (English)
Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests, while no effective mitigation yet exists. In this work, we show that this safety degradation arises from a separable component in the vector that disrupts the model's safety mechanisms but contributes little to the steering objective. We identify and remove this safety-degrading component, formulating the task as a constrained optimization problem solved through primal-dual updates, subject to preserving the intended steering effect and bounding false refusal. The resulting solution is both interpretable and surgical: the optimization recovers a single direction whose ablation from the steering vector restores model safety with minimal utility cost. Across models, steering behaviors, and attack suites, including unseen attacks types, our method substantially reduces steering-induced safety degradation while preserving the original steering effect with minimal impact on false refusal. Our method offers a post-hoc correction to steering vectors that mitigates their safety cost, and more broadly, it provides a general recipe for applying activation-level model interventions without paying a safety tax.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。