arXiv:2411.11296cs.LG2024-11被引 92

用稀疏自编码器调控激活值,提升模型拒答安全性但会损害通用能力。

Steering Language Model Refusal with Sparse Autoencoders

  • 在推理时通过增强稀疏自编码器特征来引导模型拒绝有害请求。
  • 拒答攻击防御效果提升,但多个基准任务性能普遍下降。
  • 揭示安全特征与通用能力深度耦合,适合研究模型安全机制的学者。

语言模型负责任部署需具备拒绝不当请求的能力,同时保持模型性能。现有方法多通过额外训练修改模型权重,本文探索一种新路径:在推理阶段通过放大稀疏自编码器(SAE)中与拒答相关的特征,实现对模型激活值的调控。研究发现,该方法虽能有效提升对单轮及复杂多轮越狱攻击的鲁棒性,但伴随一个此前未被充分关注的代价——在多个基准任务上出现系统性性能退化,即使在无关联的安全输入上亦然。这表明,决定拒答行为的特征可能比以往认知更深度嵌入于模型的通用语言能力之中。研究揭示了语言模型中安全相关特征的本质及其可隔离性的关键开放问题,强调在实际部署前必须理解并应对此类能力权衡机制。尽管SAE引导在提升安全性方面展现出灵活性,其应用仍需深入研究潜在代价。

原文摘要 · Abstract (English)

Responsible deployment of language models requires mechanisms for refusing unsafe prompts while preserving model performance. While most approaches modify model weights through additional training, we explore an alternative: steering model activations at inference time via amplifying sparse autoencoder (SAE) features that mediate refusal. This work uncovers a fundamental tension between SAE steering-based safety improvements and general model capabilities. While feature steering successfully improves robustness against both single-turn and challenging multi-turn jailbreak attempts, we discover that this comes at a previously underexplored cost -- systematic degradation of performance across multiple benchmark tasks, even on safe inputs with no apparent connection to refusal behavior. This suggests that features mediating refusal may be more deeply entangled with general language model capabilities than previously understood. Our findings reveal important open questions about the nature of safety-relevant features in language models and the feasibility of isolating them for targeted intervention. While SAE-based steering shows promise as a flexible approach to enhancing language model safety, our results highlight the critical need to understand and address the mechanisms behind these capability tradeoffs before such techniques can be practically deployed.

模型安全稀疏自编码器推理调控能力退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。