通过抑制有害生成特征并增强拒绝特征,提升大模型安全控制效果。
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features

- 在相同稀疏自编码器特征空间中同时抑制有害路径、增强拒绝特征
- 在GUISE等数据集上显著减少有害回复,安全拒绝率大幅提升
- 适合关注模型安全可控性的研究者与工业部署团队
基于稀疏自编码器(SAEs)的调节为无需重训练的大语言模型行为调整提供了轻量级推理时路径。通过暴露稀疏且可解释的特征,SAE调节为安全控制提供了有前景的接口,可引导有害生成走向拒绝。然而我们发现,复杂伪装提示仍可能破坏现有SAE调节方法的效果。为此,我们构建了通用隐蔽指令安全评估数据集(GUISE),系统评估此失效模式。现有单向调节方法在有害提示上无法可靠产生拒绝,表明仅增强拒绝能力在有害生成路径仍活跃时过于薄弱。因此,我们提出拒绝增强型抑制调节(REINS),在同一SAE特征空间中抑制有害生成特征并增强安全拒绝特征。在GUISE及其他数据集上的实验表明,先前方法要么干预过弱,要么通过崩溃实现表面安全,而REINS显著降低有害响应,明显提升安全拒绝率,并基本保持通用能力。
原文摘要 · Abstract (English)
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and interpretable features, SAE steering provides a promising interface for safety control that guides harmful continuations toward refusal. However, we observe that complex wrappers can still undermine existing SAE steering methods on harmful prompts. To evaluate this failure mode systematically, we construct Generalized Undercover Instruction Safety Evaluation (GUISE), a dataset of harmful prompts with complex wrappers. Existing single direction SAE steering methods do not reliably produce refusals on harmful prompts, suggesting that refusal enhancement alone can be too weak when the harmful continuation path remains active. This motivates us to propose Refusal-Enhanced INhibitory Steering (REINS), which suppresses harmful continuation features and enhances safe refusal features in the same SAE feature space. Experiments on GUISE and other datasets show that prior methods either intervene too weakly or achieve only apparent safety through collapse, while REINS substantially reduces harmful responses, markedly improves safe refusals and largely preserves general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。