arXiv:2605.20262cs.LGcs.AI2026-05

提出残差铺路法,精准控制模型拒绝行为而不影响其他功能。

Residual Paving: Diagnosing the Routing Bottleneck in Selective Refusal Editing

论文配图:Residual Paving: Diagnosing the Routing Bottleneck in Selective Refusal Editing
图 1 · 摘自论文原文
  • 分离路由选择与编辑更新,用门控机制决定是否干预
  • 编辑后拒绝率从88.6%降至4.0%,良性保留率达95.5%
  • 适合需要精准控制模型拒绝行为的研究者

我们研究选择性拒绝编辑作为三向控制问题:在指定编辑提示上诱导非拒绝行为,同时保持编辑集外的良性行为和有害拒绝。提出残差铺路法,一种用于冻结指令微调变换器的路由残差编辑方法,将路由选择(是否干预)与残差编辑能力(执行何种编辑)分离。早期层路由器预测标量门控和专家混合;激活时,提示条件下的瓶颈残差专家在后期层应用更新,保持主干不变。该分解支持一个仅替换学习到的标量门控为保留的编辑/保留标签的最优路由诊断,固定残差编辑器和冻结主干。在主要Gemma-3-4B-IT保留分割上,学习到的残差铺路法将编辑拒绝率从88.6%降低至4.0%,良性分布保留率达95.5%,有害分布保留率达87.3%。相同协议的一向引导控制在编辑成功率上弱得多,编辑拒绝率仍为86.8%(Edit-target ActAdd)和78.9%(DIM风格拒绝引导)。剩余失败表现为非目标有害保留退化:有害拒绝仍低于冻结基线,65.3%对比81.6%。在六个主干上,最优路由在所有报告行中提升保留侧诊断得分,中位增益+12.9个百分点,支持学习到的路由选择是主要观测瓶颈。在两个主干上的轨迹诊断进一步表明,模型朝向编辑目标延续方向移动,而非通用拒绝抑制。

原文摘要 · Abstract (English)

We study selective refusal editing as a three-way control problem: induce non-refusal on designated edit prompts while preserving benign behavior and harmful refusals outside the edit set. We introduce Residual Paving, a routed residual editing method for frozen instruction-tuned transformers that separates route selectivity, whether to intervene, from residual-edit capacity, what edit to apply. An early-layer router predicts a scalar gate and expert mixture; when active, prompt-conditioned bottleneck residual experts apply later-layer residual updates while leaving the backbone unchanged. This decomposition supports an oracle-routing diagnostic where only the learned scalar gate is replaced with the held-out edit/keep label, leaving the residual editor and frozen backbone fixed. On the primary Gemma-3-4B-IT held-out split, learned Residual Paving reduces edit refusal from 88.6% to 4.0%, with 95.5% benign distribution preservation and 87.3% harmful distribution preservation. Same-protocol one-direction steering controls are much weaker on edit success, leaving edit refusal at 86.8% for Edit-target ActAdd and 78.9% for DIM-style refusal steering. The remaining failure is off-target harmful-keep degradation: harmful refusal remains below the frozen-base rate, 65.3% vs. 81.6%. Across six backbones, oracle routing improves the keep-side diagnostic score on every reported row, with median gain +12.9 pp, supporting the interpretation that learned route selectivity is the main observed bottleneck. Trajectory diagnostics on two backbones further suggest directed movement toward edit-target continuations rather than generic refusal suppression.

模型编辑拒绝控制残差路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。