arXiv:2506.04250cs.LG2025-06被引 9

用可解释的激活调控,让大模型安全生成内容而不拒答。

SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs

  • 通过类别特异性向量实现精准安全控制
  • 无需成对安全数据,不降低文本质量
  • 简单有效,适合多种模型和风险场景

微调大型语言模型以适应不断变化的安全策略成本高且不切实际。机制可解释性使推理时通过隐层激活调控成为可能,但其在精确、可定制化安全调整中的潜力尚未被充分挖掘。本文提出SafeSteer方法:(i) 利用类别特异性调控向量实现更精确控制;(ii) 采用简单、无梯度的无监督方法提升安全调控效果,同时保持文本质量与主题相关性,无需显式拒答;(iii) 不依赖对比成对安全数据。我们还指出,该方法简洁高效,契合近期研究发现——简单技术在激活调控中常优于复杂方法。实验表明,该方法在多种大模型、数据集和风险类别上均有效,能实现精准控制,避免全盘拒答,并引导模型生成安全内容同时保持主题相关性。

原文摘要 · Abstract (English)

Fine-tuning large language models (LLMs) to adapt to evolving safety policies is costly and impractical. Mechanistic interpretability enables inference-time control through latent activation steering, yet its potential for precise, customizable safety adjustments remains largely untapped. This paper investigates an approach called SafeSteer for guiding the outputs of LLMs by: (i) leveraging category-specific steering vectors for more precise control, (ii) employing a simple, gradient-free unsupervised method to enhance safety steering while preserving text quality, topic relevance, and without explicit refusal, and (iii) accomplishing this without a hard requirement of contrastive pairwise safe data. We also highlight that our method, being simple and effective, aligns with recent studies suggesting that simple techniques often outperform more complex ones in activation steering. We showcase the effectiveness of our approach across various LLMs, datasets, and risk categories, demonstrating its ability to provide precise control, prevent blanket refusals, and guide models toward generating safe content while maintaining topic relevance.

大模型安全可解释性激活调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。