arXiv:2607.27081cs.AIcs.CL2026-07

提出路由式在线蒸馏方法,提升大模型安全对齐的抗模板攻击能力。

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

论文配图:On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
图 1 · 摘自论文原文
  • 通过建模对齐与污染输出分布差异,避免依赖具体提示模板。
  • 在模板不匹配时仍保持高防御效果和专业技能不退化。
  • 适合需要强鲁棒性的安全对齐场景,如内容过滤与可信AI部署。

微调是专精大语言模型(LLM)的主流范式,但存在严重漏洞:恶意数据提供者可在下游语料中嵌入有害行为,使模型在保留专业能力的同时按需违背人类价值观。现有安全对齐防御常因三大缺陷失效:易引发灾难性遗忘;当防御方无法观测攻击者提示模板时效果崩溃;成功对齐的模型仍可通过简单系统提示切换被重新越狱。为此,我们提出基于路由的在线蒸馏(ROPD),不拟合特定提示模板,而是建模对齐与污染输出概率分布的差异。我们在三个数据集和三种不同对齐强度的基线模型上,对比了ROPD与四种先进基线。结果表明,当基线方法遭遇模板不匹配时,通常伴随下游任务性能严重下降;而ROPD显著降低模板不匹配风险,在防御有效性与能力保留方面均表现更优。尽管分析显示ROPD并非完全免疫模板变化,其性能退化远小于现有方法,确立了新型鲁棒性标准。

原文摘要 · Abstract (English)

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstream task performance. In contrast, ROPD substantially mitigates template-mismatch risks, maintaining superior robustness in both defense effectiveness and capability preservation. While our analysis indicates ROPD is not entirely immune to template shifts, its performance degradation is negligible compared to existing methods, establishing a new standard for robust LLM realignment.

大模型安全对齐防御鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。