arXiv:2606.13709stat.MLcs.LG2026-06

LoMC让模型更愿回答问题,且不损失通用能力。

LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models

论文配图:LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models
图 1 · 摘自论文原文
  • 先找关键修正区域,再分层精准修正。
  • 在4个模型上显著提升非拒绝响应率。
  • 适合需要微调安全性的研究人员。

我们研究了路由型MoE及混合MoE基础模型中的可控后训练拒绝抑制问题,目标是在保持通用能力的同时,提升非拒绝目标响应行为,且干预范围紧凑。现有基于方向的广泛修改会扰动通用计算,而仅支持专家的修改又常因容量不足难以纠正异构的拒绝表征。为此,我们提出局部多向修正(LoMC),一种支持门控的干预框架:先识别紧凑的编辑支持,再将原型修正方向聚合为逐层修正方向,最后仅在选定支持范围内应用秩一逐层修正。通过将编辑支持作为结构化门控约束,LoMC在不扩大干预范围的前提下提升了修正容量。在四个路由骨干模型上的文本与多模态安全基准测试表明,LoMC显著提升了非拒绝目标响应行为,同时保持了通用能力,且干预足迹紧凑。

原文摘要 · Abstract (English)

We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving general capability under a compact intervention footprint. Existing broad direction-based edits can perturb general-purpose computation, whereas support-only expert edits often lack sufficient capacity to correct heterogeneous refusal representations. To address this limitation, we introduce Localized Multidirectional Correction (LoMC), a support-gated intervention framework that follows a support-then-correction execution order: it first identifies a compact edit support, then aggregates prototype correction directions into layer-wise correction directions, and finally applies rank-one layer-wise correction only within the selected support. By using the edit support as a structural gating constraint, LoMC increases correction capacity without expanding the intervention scope. Experiments on text-only and multimodal safety benchmarks across four routed backbones show that LoMC substantially improves non-refusal target-response behavior while maintaining general capability under a compact intervention footprint.

模型修正安全对齐MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。