arXiv:2411.07554stat.MLcs.LG2024-11被引 2

揭示外源随机性如何通过特征采样提升随机森林性能

Exogenous Randomness Empowering Random Forests

  • 引入外源随机性概念,区分特征采样与树构建中的随机类型
  • 理论证明特征采样可同时降低偏差与方差,提升模型一致性
  • 发现噪声特征因采样机制反而提升性能,适合算法优化研究者

本文从理论与实证角度分析外源随机性对独立于训练数据的树构建规则的随机森林的影响。首次正式定义外源随机性,并识别两类常见随机性:类型I来自特征子采样,类型II来自树构建过程中的平局处理。推导了个体树与森林均方误差(MSE)的非渐近展开式,确立其一致性的充分必要条件。在独立特征的线性回归模型中,MSE展开更明确,揭示随机森林机制并给出带显式一致性速率的MSE上界。基于理论结果,仿真表明特征子采样能同时降低随机森林的偏差与方差,实现偏差-方差自适应平衡;更意外的是,噪声特征因子采样机制反而成为性能提升的‘福音’。

原文摘要 · Abstract (English)

We offer theoretical and empirical insights into the impact of exogenous randomness on the effectiveness of random forests with tree-building rules independent of training data. We formally introduce the concept of exogenous randomness and identify two types of commonly existing randomness: Type I from feature subsampling, and Type II from tie-breaking in tree-building processes. We develop non-asymptotic expansions for the mean squared error (MSE) for both individual trees and forests and establish sufficient and necessary conditions for their consistency. In the special example of the linear regression model with independent features, our MSE expansions are more explicit, providing more understanding of the random forests' mechanisms. It also allows us to derive an upper bound on the MSE with explicit consistency rates for trees and forests. Guided by our theoretical findings, we conduct simulations to further explore how exogenous randomness enhances random forest performance. Our findings unveil that feature subsampling reduces both the bias and variance of random forests compared to individual trees, serving as an adaptive mechanism to balance bias and variance. Furthermore, our results reveal an intriguing phenomenon: the presence of noise features can act as a "blessing" in enhancing the performance of random forests thanks to feature subsampling.

随机森林外源随机性偏差方差权衡特征采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。