arXiv:2410.16201stat.MLcs.LG2024-10ICML被引 6

过参数化神经网络的集成模型未必比单个大模型泛化更好。

Theoretical Limitations of Ensembles in the Age of Overparameterization

  • 用随机特征回归器建模,理论证明过参数化集成等价于单个大模型。
  • 有限宽度集成迅速收敛到同等参数量的单模型,泛化能力几乎相同。
  • 集成内部预测差异反映容量增加效应,非传统不确定性度量,适合研究者参考。

经典集成模型通常比单个组件模型具有更好的泛化性能。然而,近期实证研究发现,现代过参数化神经网络的集成未必优于单个更大的神经网络。本文以随机特征(RF)回归器的集成为基础,建立理论框架,阐明现代过参数化集成与经典欠参数化集成的本质差异。在欠参数化情形下,集成通常引入正则化并提升泛化能力;但在过参数化情形下,我们以最小假设证明:无限宽度的过参数化集成等价于单一无限宽的RF回归器,有限宽度集成也快速趋近于同参数预算下的单模型。该结论对无正则项模型为严格成立,对小正则项模型为近似成立。结果表明,过参数化集成与单一大模型的泛化性能几乎一致。进一步分析显示,集成成员间的预测方差反映的是容量增加的预期影响,而非传统意义上的不确定性。这些发现挑战了过参数化场景下集成优势的普遍认知,提示需重新审视欠参数化集成直觉向深度集成和过参数化领域的迁移适用性。

原文摘要 · Abstract (English)

Classic ensembles generalize better than any single component model. In contrast, recent empirical studies find that modern ensembles of (overparameterized) neural networks may not provide any inherent generalization advantage over single but larger neural networks. This paper clarifies how modern overparameterized ensembles differ from their classic underparameterized counterparts, using ensembles of random feature (RF) regressors as a basis for developing theory. In contrast to the underparameterized regime, where ensembling typically induces regularization and increases generalization, we prove with minimal assumptions that infinite ensembles of overparameterized RF regressors become pointwise equivalent to (single) infinite-width RF regressors, and finite width ensembles rapidly converge to single models with the same parameter budget. These results, which are exact for ridgeless models and approximate for small ridge penalties, imply that overparameterized ensembles and single large models exhibit nearly identical generalization. We further characterize the predictive variance amongst ensemble members, demonstrating that it quantifies the expected effects of increasing capacity rather than capturing any conventional notion of uncertainty. Our results challenge common assumptions about the advantages of ensembles in overparameterized settings, prompting a reconsideration of how well intuitions from underparameterized ensembles transfer to deep ensembles and the overparameterized regime.

集成学习过参数化泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。