arXiv:2412.05418cs.LGcond-mat.dis-nn2024-12被引 2

固定参数量下,大模型比小模型集成更优,但过参数情形下集成可近似最优。

No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions

  • 通过随机特征岭回归分析模型规模与集成数的权衡关系
  • 参数总量不变时,集成数越多,测试误差越高;单一大模型性能最佳
  • 过参数情况下,只要总特征数相同,集成效果接近最优

在固定模型总参数量的前提下,需在训练单个大模型与组合多个小模型之间权衡。本文研究了随机特征岭回归模型在过参数与欠参数情形下的集成表现。基于确定性等价风险估计,证明当固定参数量分配给K个独立训练模型时,优化后的岭回归测试误差随K增大而上升,因此单一大模型性能最优。进一步探讨集成能否实现近似最优:在过参数情形下,测试误差仅依赖于总特征数,故过参数集成始终接近最优。对于欠参数情形,推导出在联合按生长指数ℓ缩放模型数量与每成员参数量时的测试误差标度律。尽管最优标度始终由固定集成数增大模型规模实现,但本研究揭示了在核函数与任务特征结构满足特定条件下,联合缩放也能获得近似最优标度。

原文摘要 · Abstract (English)

Given a fixed budget for total model size, one must choose between training a single large model or combining the predictions of multiple smaller models. We investigate this trade-off for ensembles of random-feature ridge regression models in both the overparameterized and underparameterized regimes. Using deterministic equivalent risk estimates, we prove that when a fixed number of parameters is distributed among $K$ independently trained models, the ridge-optimized test risk increases with $K$. Consequently, a single large model achieves optimal performance. We then ask when ensembles can achieve \textit{near}-optimal performance. In the overparameterized regime, we show that, to leading order, the test error depends on ensemble size and model size only through the total feature count, so that overparameterized ensembles consistently achieve near-optimal performance. To understand underparameterized ensembles, we derive scaling laws for the test risk as a function of total parameter count when the ensemble size and parameters per ensemble member are jointly scaled according to a ``growth exponent'' $\ell$. While the optimal error scaling is always achieved by increasing model size with a fixed ensemble size, our analysis identifies conditions on the kernel and task eigenstructure under which near-optimal scaling laws can be obtained by joint scaling of ensemble size and model size.

集成学习随机特征标度律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。