揭示子采样中默认比例为何失效,并提出动态调整策略
Exact Finite-Sample Variance Decomposition of Subagging: A Spectral Filtering Perspective
- 用谱滤波视角精确分解子采样的方差,无需渐近假设
- 发现子采样本质是低通滤波,高阶交互方差按α^c衰减
- 提出根据学习器复杂度自适应调整采样率,提升泛化性能
标准子采样比例(如 α≈0.632)在集成学习中作为默认基准已使用三十年。然而,这些比例如何与基学习器的内在函数复杂度在有限样本下相互作用,尚无确切数学刻画。本文利用 Hoeffding-ANOVA 分解,首次推导出适用于任意对称基学习器、无需渐近极限或光滑性假设的子采样精确有限样本方差分解。我们证明子采样相当于确定性低通谱滤波:保留低阶结构信号,同时将第 c 阶交互方差按几何因子趋近于 α^c 的速率衰减。该解耦机制揭示了为何默认比例常对高容量插值器欠正则化,后者需更小的 α 以指数抑制虚假高阶噪声。为实现实用化,我们提出一种基于复杂度引导的自适应子采样算法,实验表明动态校准 α 至学习器复杂度谱能持续优于静态基准。
原文摘要 · Abstract (English)
Standard resampling ratios (e.g., $α\approx 0.632$) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base learner's intrinsic functional complexity in finite samples lacks a exact mathematical characterization. We leverage the Hoeffding-ANOVA decomposition to derive the first exact, finite-sample variance decomposition for subagging, applicable to any symmetric base learner without requiring asymptotic limits or smoothness assumptions. We establish that subagging operates as a deterministic low-pass spectral filter: it preserves low-order structural signals while attenuating $c$-th order interaction variance by a geometric factor approaching $α^c$. This decoupling reveals why default baselines often under-regularize high-capacity interpolators, which instead require smaller $α$ to exponentially suppress spurious high-order noise. To operationalize these insights, we propose a complexity-guided adaptive subsampling algorithm, empirically demonstrating that dynamically calibrating $α$ to the learner's complexity spectrum consistently improves generalization over static baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。