揭示正则化如何改变模型训练中的隐式偏置方向与范围
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
- 将正则化纳入镜像流框架,分析其对训练轨迹几何的影响
- 发现正则化可调控隐式偏置的位置、类型和范围收缩
- 实验证明动态关闭权重衰减能提升泛化性能
隐式偏置在解释过参数化模型为何能良好泛化方面起着关键作用。实践中常结合显式正则化(如权重衰减)以防止过拟合。尽管两者常被分别研究,但实际中往往协同作用。理解它们的相互作用,是控制隐式偏置的形状与强度的关键,因为显式正则化可改变其特性。为此,本文将显式正则化引入镜像流框架,分析其对训练动态几何的长期影响,涵盖三类效应:位置偏置、偏置类型及范围收缩。我们的分析覆盖稀疏编码、矩阵感知、单层注意力和LoRA等多类问题,验证了理论洞察的实用性。为进一步利用正则化的长期效应,我们提出在训练中切换关闭权重衰减,实验表明该策略可提升模型泛化能力。
原文摘要 · Abstract (English)
Implicit bias plays an important role in explaining how overparameterized models generalize well. Explicit regularization like weight decay is often employed in addition to prevent overfitting. While both concepts have been studied separately, in practice, they often act in tandem. Understanding their interplay is key to controlling the shape and strength of implicit bias, as it can be modified by explicit regularization. To this end, we incorporate explicit regularization into the mirror flow framework and analyze its lasting effects on the geometry of the training dynamics, covering three distinct effects: positional bias, type of bias, and range shrinking. Our analytical approach encompasses a broad class of problems, including sparse coding, matrix sensing, single-layer attention, and LoRA, for which we demonstrate the utility of our insights. To exploit the lasting effect of regularization and highlight the potential benefit of dynamic weight decay schedules, we propose to switch off weight decay during training, which can improve generalization, as we demonstrate in experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。