揭示宽度与数据量如何共同决定模型泛化性能
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

- 在有限样本下分析二次网络的泛化误差,明确参数量与样本数关系
- 发现泛化误差服从由目标谱结构决定的幂律,存在不同阶段的相变
- 适合研究模型缩放规律、泛化理论的学者参考
理解性能随模型规模和数据量协同变化是现代机器学习的核心问题。现有理论通常将泛化描述为数据或计算量的函数,常在固定特征或无限宽度假设下进行,且针对在线SGD。本文则研究特征学习模型中可训练参数数量与样本数对泛化的影响。我们在有限样本设置下,分析二次两层网络中ℓ₂正则化的经验测试误差最小化问题,利用结构化数据实现对泛化误差随样本数、模型宽度和正则化程度的显式刻画。结果揭示了随着参数量变化的相图,呈现不同缩放阶段。特别地,泛化误差遵循由目标函数谱结构决定的数据相关幂律。我们进一步刻画了各阶段间的转变,包括插值开始出现的临界点及其对泛化的影响。
原文摘要 · Abstract (English)
Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical works on scaling laws typically describe generalization as a function of data or compute, often in fixed-feature or infinite-width regimes and for online SGD. Here, we instead study how generalization scales with the number of trainable parameters and the number of samples in a feature-learning model. We analyze $\ell_2$-regularized empirical test error minimization in a quadratic two-layer network in a finite-sample setting with structured data. This setting allows for an explicit characterization of the generalization error as a function of the number of samples, model width, and regularization. Our results reveal a phase diagram with distinct scaling regimes as the number of parameters varies. In particular, the generalization error follows data-dependent power laws controlled by the spectral structure of the target. We further characterize the transitions between regimes, including the onset of interpolation, and their impact on generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。