arXiv:2608.02674cs.CRcs.AI2026-08

动态路由自适应对齐,让大模型在安全路径被攻破时仍能拒绝有害请求。

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

论文配图:Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
图 1 · 摘自论文原文
  • 通过对比安全与不安全样本激活模式,定位模型内部安全路径。
  • 主动屏蔽安全路径生成对抗样本,构建故障感知的偏好数据对。
  • 动态生成补偿路由,显著提升对白盒攻击的鲁棒性,适合高安全需求场景。

随着大型基础模型(LFMs)在开放环境中的广泛应用,安全威胁正从黑盒越狱转向直接识别并破坏内部安全神经元或路径的白盒攻击。现有安全防御多依赖静态安全单元或固定拒答路径,难以应对针对性的路径级白盒攻击。为此,我们提出动态路由自适应对齐(DRAA)框架,引入动态补偿路径以在安全路径被破坏时仍保持稳健拒答行为。首先,通过对比安全与不安全校准样本的内部激活,定位模型的安全路径。DRAA随后屏蔽该路径,诱导因果失败案例,并选择性挖掘由此产生的防御失败,进而构建故障感知的偏好对。大量实验表明,DRAA有效重构了模型安全的底层路径依赖,显著提升了对路径级白盒攻击的鲁棒性,同时保持了通用能力。

原文摘要 · Abstract (English)

With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, existing safety defenses often rely on static safety units or fixed refusal pathways, leaving models highly vulnerable to targeted route-level white-box attacks. For that, we propose dynamic routing adaptive alignment (DRAA), a framework that introduces dynamic compensatory routes to preserve robust refusal behavior when the safety route is compromised. Specifically, we first identify and localize the model's safety route by contrasting internal activations between safe and unsafe calibration samples. DRAA then masks this safety route to induce causal failure cases and selectively mines the resulting defense failures, thereby constructing failure-aware preference pairs. Extensive experiments demonstrate that DRAA effectively restructures the underlying pathway dependence of model safety, substantially improving robustness against route-level white-box attacks, while preserving general utility.

安全防御白盒攻击动态路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。