提出改进的随机森林方法,解决数据不平衡问题并降低预测方差。
Asymptotic Normality of Infinite Centered Random Forests -Application to Imbalanced Classification
- 用重要性采样重构数据,再对中心化随机森林去偏。
- 在极端不平衡场景下,新方法方差显著低于原始方法。
- 理论证明适用于经典随机森林,实验验证效果稳定。
许多分类任务面临数据不平衡问题,即某一类别严重样本不足。现有方法通常通过构建平衡数据集来训练分类器。本文理论研究了当分类器为中心化随机森林(CRF)时,此类数据重平衡策略的效果。建立了无限规模CRF的中心极限定理(CLT),给出明确收敛速率与精确常数。进一步证明,在平衡数据上训练的CRF存在偏差,可通过适当技术消除。基于重要性采样(IS)方法,提出去偏估计器IS-ICRF,其满足以真实预测函数值为中心的CLT。在高度不平衡情形下,理论与实验均表明,IS-ICRF的方差低于直接在原始数据上训练的ICRF。理论结果特别是方差率与方差缩减效应,在实验中也适用于Breiman随机森林。
原文摘要 · Abstract (English)
Many classification tasks involve imbalanced data, in which a class is largely underrepresented. Several techniques consists in creating a rebalanced dataset on which a classifier is trained. In this paper, we study theoretically such a procedure, when the classifier is a Centered Random Forests (CRF). We establish a Central Limit Theorem (CLT) on the infinite CRF with explicit rates and exact constant. We then prove that the CRF trained on the rebalanced dataset exhibits a bias, which can be removed with appropriate techniques. Based on an importance sampling (IS) approach, the resulting debiased estimator, called IS-ICRF, satisfies a CLT centered at the prediction function value. For high imbalance settings, we prove that the IS-ICRF estimator enjoys a variance reduction compared to the ICRF trained on the original data. Therefore, our theoretical analysis highlights the benefits of training random forests on a rebalanced dataset (followed by a debiasing procedure) compared to using the original data. Our theoretical results, especially the variance rates and the variance reduction, appear to be valid for Breiman's random forests in our experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。