arXiv:2602.04448cs.LGcs.AI2026-02被引 2

针对专家模型的安全对齐难题,提出路由感知的精准修复方法

RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models

  • 识别被越狱攻击频繁激活的敏感专家,仅对其精细微调
  • 在固定路由下修复关键专家,实现近似完美抗攻击能力
  • 适合关注MoE模型安全性的研究者与工业应用开发者

混合专家(MoE)语言模型因稀疏路由机制带来独特的安全对齐挑战,标准全参数微调可能因路由或专家主导效应降低攻击成功率,而非直接修复安全关键专家。本文提出RASA框架,通过识别被成功越狱攻击频繁激活的专家,仅在固定路由下选择性微调这些专家,并随后强制路由与安全上下文一致。在两种代表性MoE架构和多种越狱攻击测试中,RASA实现近乎完美的鲁棒性、强跨攻击泛化能力,显著减少过度拒绝,同时保持MMLU、GSM8K和TruthfulQA等基准上的通用能力。结果表明,针对特定专家的精准修复优于全局参数更新,为安全对齐提供了实用且保留架构的替代方案。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.

MoE模型安全对齐专家修复越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。