arXiv:2410.01259stat.MLcs.LG2024-10被引 2

重新定义模型复杂度,揭示过参数化模型为何仍能泛化。

Revisiting Optimism and Model Complexity in the Wake of Overparameterized Machine Learning

  • 从有效自由度出发,提出适配随机输入的复杂度度量
  • 在过参数化下,模型仍可实现良好泛化性能
  • 适用于分析任意预测模型的泛化能力

现代机器学习普遍采用参数远超样本数量的过参数化模型,这类模型表现出令人意外的泛化能力,例如预测误差随参数量变化呈现‘双下降’现象。本文从基础出发,重新诠释并扩展经典统计学中的(有效)自由度概念。传统定义关联固定协变量下的预测误差(基于训练时相同的非随机协变量点),而本文提出的扩展定义则关联随机协变量下的预测误差(基于从协变量分布中抽取的新随机样本)。这一随机-协变量设定更贴合现代机器学习场景:即使模型复杂到能插值训练数据,仍可在适当条件下实现良好泛化。我们通过概念论证、理论推导与实验验证了所提复杂度度量的有效性,并展示了其在解释与比较任意预测模型中的应用潜力。

原文摘要 · Abstract (English)

Common practice in modern machine learning involves fitting a large number of parameters relative to the number of observations. These overparameterized models can exhibit surprising generalization behavior, e.g., ``double descent'' in the prediction error curve when plotted against the raw number of model parameters, or another simplistic notion of complexity. In this paper, we revisit model complexity from first principles, by first reinterpreting and then extending the classical statistical concept of (effective) degrees of freedom. Whereas the classical definition is connected to fixed-X prediction error (in which prediction error is defined by averaging over the same, nonrandom covariate points as those used during training), our extension of degrees of freedom is connected to random-X prediction error (in which prediction error is averaged over a new, random sample from the covariate distribution). The random-X setting more naturally embodies modern machine learning problems, where highly complex models, even those complex enough to interpolate the training data, can still lead to desirable generalization performance under appropriate conditions. We demonstrate the utility of our proposed complexity measures through a mix of conceptual arguments, theory, and experiments, and illustrate how they can be used to interpret and compare arbitrary prediction models.

模型复杂度过参数化泛化能力统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。