arXiv:2409.14123stat.MLcs.LG2024-09

解释大模型泛化能力为何随规模增长趋于稳定

Consistency for Large Neural Networks: Regression and Classification

  • 证明参数越多,逼近误差越小但泛化误差被正则约束
  • 测试误差最终收敛到由泛化与优化误差决定的常数
  • 为回归与分类任务提供理论依据,适合研究泛化者

尽管过参数化模型在实践中取得显著成功,其理论性质尤其是泛化行为仍不完全清楚。众所周知的双下降现象表明,神经网络的测试误差随模型规模增大单调下降,最终收敛至非零常数。本文旨在揭示这一尾部行为的理论机制,并研究深度过参数化神经网络在多种学习任务(包括回归和分类)中的统计一致性。首先,我们证明当参数数量增加时,逼近误差单调减小,而显式或隐式正则化(如权重衰减)使泛化误差保持存在但有界。因此,整体误差曲线最终收敛至由有界泛化误差和优化误差决定的常数。其次,我们证明若使用正则化技术,深度过参数化神经网络在多个学习任务中具有统计一致性。理论结果与数值实验一致,为理解过参数化神经网络的泛化行为提供了新视角。

原文摘要 · Abstract (English)

Although overparameterized models have achieved remarkable practical success, their theoretical properties, particularly their generalization behavior, remain incompletely understood. The well known double descents phenomenon suggests that the test error curve of neural networks decreases monotonically as model size grows and eventually converges to a non-zero constant. This work aims to explain the theoretical mechanism underlying this tail behavior and study the statistical consistency of deep overparameterized neural networks in many different learning tasks including regression and classification. Firstly, we prove that as the number of parameters increases, the approximation error decreases monotonically, while explicit or implicit regularization (e.g., weight decay) keeps the generalization error existing but bounded. Consequently, the overall error curve eventually converges to a constant determined by the bounded generalization error and the optimization error. Secondly, we prove that deep overparameterized neural networks are statistical consistency across multiple learning tasks if regularization technique is used. Our theoretical findings coincide with numerical experiments and provide a perspective for understanding the generalization behavior of overparameterized neural networks.

泛化理论过参数化神经网络统计一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。