arXiv:2607.26964stat.MLcs.LG2026-07

通过特征不稳定度分析,证明特征袋装能显著提升模型稳定性。

Feature Bagging Provides Stability

  • 引入特征不稳定度衡量移除单个特征的影响。
  • 袋装后稳定性显著提升,尤其在激进子采样时效果更明显。
  • 少量袋装轮次即可接近无限袋装的稳定水平,适合实际应用。

我们从算法稳定性的角度研究特征袋装。特征袋装是一种集成策略,通过在随机子采样的特征子集上训练基学习器并聚合结果,可能依赖数据。我们提出特征不稳定度(FI),作为实例不稳定度(II)在特征维度上的类比,用于衡量模型对移除单个特征的敏感性。较小的FI或II值对应更强的稳定性,实验表明FI能捕捉与泛化性能相关的互补信息。在此框架下,我们在参数线性模型和受随机森林递归特征子采样启发的模型无关设置中分析了特征袋装。两种情形下均建立了形式化保证:特征袋装相比非袋装版本提升了相关稳定性,且在更激进的子采样条件下提升更大。进一步证明,少量袋装轮次即足以逼近无限袋装的稳定性水平。

原文摘要 · Abstract (English)

We study feature bagging through the lens of algorithmic stability. Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner. We introduce feature instability (FI), the feature-axis analogue of instance instability (II), which measures sensitivity to removing a single feature. Smaller values of II or FI correspond to stronger stability, and our experiments show that FI captures generalization-relevant information complementary to II. Within this framework, we analyze feature bagging in both a parametric linear model and a model-free setting inspired by recursive feature subsampling in random forests. In both settings, we establish formal guarantees showing that feature bagging improves the relevant stability relative to its non-bagged counterpart, with larger improvements under more aggressive subsampling. We further show that a modest number of bagging rounds is sufficient to approach the infinite-bagging stability level.

集成学习稳定性分析特征选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。