用范数控制容量,揭示深度模型泛化规律
The $φ$ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control
- 基于随机特征估计器,用模型范数替代参数量衡量复杂度
- 发现学习曲线存在从欠拟合到过拟合的相变,无双下降现象
- 为理解大模型泛化提供新视角,适合研究理论机制的读者
理解测试风险随模型复杂度的变化是机器学习的核心问题。经典理论难以解释大规模过参数化深度网络的观测学习曲线。基于参数数量的容量度量通常无法解释这些现象。为此,我们采用基于范数的容量度量,并在广泛用于简化理论分析的随机特征估计器框架下开展研究。结果表明,估计器的范数会集中,且其大小决定测试误差。预测的学习曲线表现出从欠参数化到过参数化的相变,但不存在双下降行为。这证实:当使用基于模型范数而非规模的容量度量时,更经典的U形泛化曲线得以恢复。技术上,我们以确定性等价为核心工具,进一步提出了新的确定性量,具有独立研究价值。
原文摘要 · Abstract (English)
Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learning curves observed for large over-parametrized deep networks. Capacity measures based on parameter count typically fail to account for these empirical observations. To tackle this challenge, we consider norm-based capacity measures and develop our study for random features based estimators, widely used as simplified theoretical models for more complex networks. In this context, we provide a precise characterization of how the estimator's norm concentrates and how it governs the associated test error. Our results show that the predicted learning curve admits a phase transition from under- to over-parameterization, but no double descent behavior. This confirms that more classical U-shaped behavior is recovered considering appropriate capacity measures based on models norms rather than size. From a technical point of view, we leverage deterministic equivalence as the key tool and further develop new deterministic quantities which are of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。