提出结构化加权理论,解释为何稳定模型的集成仍有效
A General Weighting Theory for Ensemble Learning: Beyond Variance Reduction via Spectral and Geometric Structure
- 将集成学习视为假设空间上的线性算子,引入谱与几何约束
- 非均匀加权可重塑逼近几何,即使方差降低不明显也优于平均
- 涵盖经典平均、堆叠及斐波那契加权,适用于平滑模型
传统集成学习被解释为方差缩减策略,适用于决策树等不稳定预测器。然而,该解释无法说明光滑样条、核岭回归、高斯过程回归等内在稳定估计器的集成效果,这些方法因正则化和谱收缩已具备严格控制的方差。本文建立了一种通用加权理论,突破经典方差缩减框架。我们将集成视为作用于假设空间的线性算子,并在加权序列空间中引入几何与谱约束。在此框架下,推导出精细的偏差-方差近似分解,表明非均匀、结构化的权重可通过重塑逼近几何与重分配谱复杂度,超越均匀平均,即使方差降低可忽略。主要结果给出了结构加权严格优于均匀集成的条件,并表明最优权重是约束二次规划的解。经典平均、堆叠及最近提出的斐波那契集成均为此统一理论的特例,该理论还兼容几何型、亚指数型与重尾加权律。整体工作为结构驱动的集成学习提供了原则性基础,解释了为何集成对平滑、低方差基学习器依然有效,并为后续发展分布自适应与动态演化加权方案奠定基础。
原文摘要 · Abstract (English)
Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. Within this framework, we derive a refined bias-variance approximation decomposition showing how non-uniform, structured weights can outperform uniform averaging by reshaping approximation geometry and redistributing spectral complexity, even when variance reduction is negligible. Our main results provide conditions under which structured weighting provably dominates uniform ensembles, and show that optimal weights arise as solutions to constrained quadratic programs. Classical averaging, stacking, and recently proposed Fibonacci-based ensembles appear as special cases of this unified theory, which further accommodates geometric, sub-exponential, and heavy-tailed weighting laws. Overall, the work establishes a principled foundation for structure-driven ensemble learning, explaining why ensembles remain effective for smooth, low-variance base learners and setting the stage for distribution-adaptive and dynamically evolving weighting schemes developed in subsequent work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。