arXiv:2604.24350cs.LGcs.AI2026-04

将灾难性过拟合解释为隐蔽后门机制,提出针对性缓解方法。

Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training

论文配图:Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
图 1 · 摘自论文原文
  • 从后门视角重新解读灾难性过拟合,揭示其本质为弱触发的不可学习任务。
  • 实验验证在多种攻击下模型仍存在泛化失败,且可被特定触发器激活。
  • 基于后门防御策略,通过参数重校准与权重异常抑制有效缓解过拟合。

快速对抗训练(FAT)因其提升神经网络对对抗攻击鲁棒性的高效性而受到广泛关注。然而,FAT易发生灾难性过拟合(CO),即模型过度拟合训练中使用的特定攻击,无法泛化到其他攻击。尽管已有方法提出多种假设并尝试缓解CO,但尚缺乏系统且直观的解释。本文创新性地从后门角度解析CO:通过路径划分、特征预测多样性及通用类别可区分触发器的验证,将CO视为不可学习任务的一种弱触发变体,统一了CO、后门攻击与不可学习任务的理论框架。基于此,提出两种受后门启发的缓解策略:(i) 使用普通微调、线性探测或重初始化技术重校准受影响的模型参数;(ii) 引入权重异常抑制约束,调控模型权重中的异常偏移。大量实验验证了该解释的合理性,并证明所提方法的有效性。

原文摘要 · Abstract (English)

Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone to catastrophic overfitting (CO), wherein models overfit to the specific attack used during training and fail to generalize to others. While existing methods introduce diverse hypotheses and propose various strategies to mitigate CO, a systematic and intuitive explanation of CO remains absent. In this work, we innovatively interpret CO through the lens of backdoor. Through validations on pathway division, diverse feature predictions, and universal class distinguishable triggers in CO, we conceptualize CO as a weak trigger variant of unlearnable tasks, unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework. Guided by this, we leverage several backdoor inspired strategies to mitigate CO: (i) Recalibrate CO affected model parameters using vanilla fine tuning, linear probing, or reinitialization-based techniques; (ii) Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies.

对抗训练后门攻击过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。