提出高维持续学习中L2正则的最优强度与任务数的近线性关系。
Optimal L2 Regularization in High-dimensional Continual Linear Regression
- 用闭式公式推导高维下持续线性回归的泛化损失。
- 最优正则强度随任务数T近似线性增长,为T/ln T。
- 首次理论揭示持续学习正则化最佳强度,适合系统设计者。
我们研究了过参数化持续线性回归中的泛化问题,模型在一系列任务上使用L2(各向同性)正则化进行训练。在高维极限下,我们推导出适用于任意线性教师的期望泛化损失的闭式表达。结果表明,在单教师和多个独立同分布教师设置下,各向同性正则化能有效缓解标签噪声,而此前处理多教师的方法要么未使用正则化,要么需要高内存开销。进一步证明,最优固定正则化强度几乎随任务数T线性增长,具体为T/ln T。据我们所知,这是首个关于理论持续学习中正则化强度的此类结果。最后,通过在线性回归和神经网络上的实验验证了理论发现,展示了该缩放规律对泛化的影响,并为持续学习系统的设计提供了实用指导。
原文摘要 · Abstract (English)
We study generalization in an overparameterized continual linear regression setting, where a model is trained with L2 (isotropic) regularization across a sequence of tasks. We derive a closed-form expression for the expected generalization loss in the high-dimensional regime that holds for arbitrary linear teachers. We demonstrate that isotropic regularization mitigates label noise under both single-teacher and multiple i.i.d. teacher settings, whereas prior work accommodating multiple teachers either did not employ regularization or used memory-demanding methods. Furthermore, we prove that the optimal fixed regularization strength scales nearly linearly with the number of tasks $T$, specifically as $T/\ln T$. To our knowledge, this is the first such result in theoretical continual learning. Finally, we validate our theoretical findings through experiments on linear regression and neural networks, illustrating how this scaling law affects generalization and offering a practical recipe for the design of continual learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。