arXiv:2601.08489cs.CL2026-01被引 1

通过分离拒绝信号与模型能力,实现安全拒绝对话的精准调控。

Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning

  • 构建概念原子库,用谱正则化方法剥离拒绝向量中的干扰成分
  • 在五种模型上实现0-2%拒绝率,困惑度变化仅0.02(Wikitext-2)
  • 适合需要保持模型能力又强化安全性的大模型定制场景

安全对齐的语言模型会系统性拒绝有害请求。虽然激活操控可调节拒绝行为,但直接删去从对比有害与无害提示中计算出的原始‘拒绝向量’常引发副作用和分布漂移。我们指出,这种退化源于原始向量的多义性,其将拒绝信号与核心能力电路及语言风格混杂。为此提出外科式拒绝消融(SRA),通过构建独立的概念原子库(代表受保护能力与风格干扰),利用岭正则化谱残差化方法,使拒绝向量正交于这些方向。所得清洁的拒绝方向仅作用于拒绝相关结构,最小化对模型语义几何的破坏。在五个模型(Qwen3-VL 和 Ministral 系列)上,SRA 实现深度拒绝降低(0-2%),对 Wikitext-2 的困惑度影响极小(平均ΔPPL ≈ 0.02),分布漂移微弱。值得注意的是,标准消融在 Qwen3-VL-4B 上导致严重漂移(首词KL = 2.088),而 SRA 在维持相同0%拒绝率的同时,保持原分布(KL = 0.044)。通过教师强制困惑度在 GSM8K 与 MBPP 上作为高分辨率能力代理,结果表明 SRA 有效保留数学与代码分布。这表明,常见‘模型损伤’实为‘幽灵噪声’——即脏的拒绝方向向能力子空间的谱泄漏。

原文摘要 · Abstract (English)

Safety-aligned language models systematically refuse harmful requests. While activation steering can modulate refusal, ablating the raw "refusal vector" calculated from contrastive harmful and harmless prompts often causes collateral damage and distribution drift. We argue this degradation occurs because the raw vector is polysemantic, entangling the refusal signal with core capability circuits and linguistic style. We introduce Surgical Refusal Ablation (SRA) to distill these steering directions. SRA constructs a registry of independent Concept Atoms representing protected capabilities and stylistic confounds, then uses ridge-regularized spectral residualization to orthogonalize the refusal vector against these directions. This yields a clean refusal direction that targets refusal-relevant structure while minimizing disruption to the model's semantic geometry. Across five models (Qwen3-VL and Ministral series), SRA achieves deep refusal reduction (0-2%) with negligible perplexity impact on Wikitext-2 (mean delta PPL approx. 0.02) and minimal distribution drift. Notably, standard ablation on Qwen3-VL-4B induces severe drift (first-token KL = 2.088), whereas SRA maintains the original distribution (KL = 0.044) while achieving the same 0% refusal rate. Using teacher-forced perplexity on GSM8K and MBPP as a high-resolution capability proxy, we show SRA preserves math and code distributions. These results suggest that common "model damage" is often "Ghost Noise," defined as the spectral bleeding of the dirty refusal direction into capability subspaces.

模型安全拒绝控制能力保持谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。