揭示对抗训练中稳健过拟合的动态机制,解释为何调参会影响模型泛化能力。
How Learning Dynamics Drive Adversarially Robust Generalization?
- 将对抗训练建模为带动量的离散动力系统,用概率框架分析泛化边界
- 发现学习率、损失曲率和小批量梯度共同影响模型的稳健泛化性能
- 提出通过权重扰动抑制主导曲率模式来缓解过拟合,但惩罚过度反而不利
尽管对抗训练被广泛视为构建鲁棒模型的标准范式,却仍面临稳健过拟合问题。现有经验与理论研究未能提供令人满意的机制解释。本文将带有动量的随机梯度下降(momentum SGD)下的对抗训练建模为离散时间动力系统,提出一个基于 PAC-Bayesian 的分析框架,推导出时间解析的稳健泛化界。该框架可追踪后验均值与协方差在稳态与非稳态瞬态下的闭式演化过程,揭示模型稳健泛化性能与学习率、局部损失几何结构及小批量随机梯度之间的关联。通过估计界中的关键量,阐明了稳健过拟合的内在机制。此外,框架表明对抗权重扰动可通过抑制主导损失曲率模式来减小稳健泛化差距,但过度惩罚可能对优化造成负面影响。
原文摘要 · Abstract (English)
Despite being widely adopted as a canonical framework for learning robust models, adversarial training suffers from robust overfitting. Existing empirical and theoretical explorations fail to provide a satisfactory mechanistic interpretation of the phenomenon. By modeling adversarial training with momentum SGD as a discrete-time dynamical system, we propose a PAC-Bayesian analytical framework that proves time-resolved robust generalization bounds. Specifically, our framework tracks the closed-form evolution of the posterior mean and covariance under both stationary and non-stationary transient regimes, connecting the model's robust generalization performance to learning rate, local loss geometry, and mini-batch stochastic gradients. By estimating the key quantities associated with the bound, we illustrate the underlying mechanism of robust overfitting. Our framework also shows how adversarial weight perturbation reduces robust generalization gaps by suppressing dominant loss-curvature modes, while suggesting that excessive penalization can be sub-optimal for optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。