arXiv:2509.22067cs.LGcs.AI2025-09被引 15

激活引导技术可能破坏大模型安全,让其更容易响应有害请求。

The Rogue Scalpel: Activation Steering Compromises LLM Safety

  • 通过向隐藏状态注入语义向量操控模型行为
  • 随机方向引导使有害响应概率升至1%-13%
  • 通用攻击可跨任务触发模型违规,适合安全研究者关注

激活引导是一种在推理时直接向模型隐藏状态注入语义向量以控制大语言模型行为的潜力技术,常被视为比微调更精确、可解释且更安全。我们证明相反情况:引导会系统性破坏模型对齐保护机制,使其更易响应有害请求。在多个模型族上进行的大量实验显示,即使在随机方向上引导,有害响应概率也能从0%上升至1%-13%。令人担忧的是,来自稀疏自编码器(SAE)的良性特征方向也表现出相当的有害潜力。最后,我们将20个随机采样的向量组合,成功构建出可破解单一提示的通用攻击,显著提升模型在未见请求上的有害响应率。这些结果挑战了‘可解释即安全’的范式,表明对模型内部的精准控制并不能保证对其行为的精准控制。

原文摘要 · Abstract (English)

Activation steering is a promising technique for controlling LLM behavior by adding semantically meaningful vectors directly into a model's hidden states during inference. It is often framed as a precise, interpretable, and potentially safer alternative to fine-tuning. We demonstrate the opposite: steering systematically breaks model alignment safeguards, making it comply with harmful requests. Through extensive experiments on different model families, we show that even steering in a random direction can increase the probability of harmful compliance from 0% to 1-13%. Alarmingly, steering benign features from a sparse autoencoder (SAE), a common source of interpretable directions, demonstrates a comparable harmful potential. Finally, we show that combining 20 randomly sampled vectors that jailbreak a single prompt creates a universal attack, significantly increasing harmful compliance on unseen requests. These results challenge the paradigm of safety through interpretability, showing that precise control over model internals does not guarantee precise control over model behavior.

模型安全激活引导对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。