arXiv:2606.05958cs.LG2026-06

操纵激活向量可无声实施模型越狱攻击

Steering Vectors are an Adversarial Attack Surface

  • 用4%-6%的恶意令牌污染数据集,生成对抗性向量
  • 攻击成功率20%-55%,较纯净向量提升19%-51%
  • 防御方法可恢复82%性能损失且不影响正常使用

激活操控已成为无需微调即可控制大语言模型行为的流行方法。由于该技术即插即用,用户常共享数据集和预计算的向量以实现行为调节。然而,我们揭示了一种隐蔽的数据投毒攻击,能无声破坏这一流程。通过替换4%-6%的训练数据中的词元,攻击者可使生成的向量对齐反拒绝方向,从而实现模型越狱,同时保持对良性提示的原有意图。在此威胁模型下,恶意实体可分发看似安全的包含文本、向量与权重的组合包,并附带可验证的等价性证书。我们在两个开源模型家族及八种模型-属性组合上测试该攻击,发现中毒向量的绝对攻击成功率(ASR)达20%-55%,较干净参照组高出19%-51%。最后,我们发现正交化拒绝方向的防御机制可在不损害良性行为的前提下,恢复约82%的攻击成功率差距。

原文摘要 · Abstract (English)

Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning. Since the technique is plug-and-play, users share datasets and precomputed vectors to steer model activations. However, we show that a \emph{stealth data poisoning attack} silently compromises this pipeline. By substituting $4{-}6\%$ of tokens in the steering dataset, an attacker can silently align the resulting vector with an anti-refusal direction. This jailbreaks the target model while preserving the intended steering effect on benign prompts. Under this threat model, a malicious actor can distribute an apparently safe bundle containing texts, vectors, and weights, alongside an equivalence certificate that the end-user can verify. We test the attack on two open-weight model families and eight model-attribute combinations, observing that poisoned vectors reach an absolute attack success rate (ASR) of $20{-}55\%$, $+19\%$ to $+51\%$ over a clean reference. Finally, we find that a refusal-direction orthogonalization defense can recover ${\approx}82\%$ of the ASR gap without harming benign behavior.

模型安全对抗攻击向量操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。