发现操控大模型行为的向量会显著影响其安全防护能力
Analysing the Safety Pitfalls of Steering Vectors
- 用对比激活添加法生成操控向量,系统测试其安全影响
- 向量可使攻击成功率最高升57%或降50%,取决于方向
- 安全漏洞源于操控向量与拒绝行为隐空间重叠
激活操控已成为无需更新权重即可调整大语言模型行为的强大工具。尽管其脆弱性和不可靠性已被广泛记录,但其安全影响仍缺乏研究。本文针对广泛应用的对比激活添加(CAA)方法生成的操控向量,采用统一评估协议进行系统性安全审计。以JailbreakBench为基准,结果显示操控向量会持续影响越狱攻击的成功率,在简单模板攻击下放大效应更明显。在不同模型家族和规模中,特定方向的操控可使攻击成功率最高提升57%或降低50%,具体取决于目标行为。我们归因于操控向量与拒绝行为潜在方向的重叠,提供了可追溯的解释。研究揭示了大模型中此前未被注意的安全缺口,凸显可控性与安全性之间的权衡。
原文摘要 · Abstract (English)
Activation steering has emerged as a powerful tool to shape LLM behavior without the need for weight updates. While its inherent brittleness and unreliability are well-documented, its safety implications remain underexplored. In this work, we present a systematic safety audit of steering vectors obtained with Contrastive Activation Addition (CAA), a widely used steering approach, under a unified evaluation protocol. Using JailbreakBench as benchmark, we show that steering vectors consistently influence the success rate of jailbreak attacks, with stronger amplification under simple template-based attacks. Across LLM families and sizes, steering the model in specific directions can drastically increase (up to 57%) or decrease (up to 50%) its attack success rate (ASR), depending on the targeted behavior. We attribute this phenomenon to the overlap between the steering vectors and the latent directions of refusal behavior. Thus, we offer a traceable explanation for this discovery. Together, our findings reveal the previously unobserved origin of this safety gap in LLMs, highlighting a trade-off between controllability and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。