为随机森林等集成方法提供统计不确定性量化的新方法
Empirical Likelihood for Random Forests and Ensembles
- 基于不完整U统计量结构构建经验似然统计量
- 修正后方法在稀疏采样下仍保持准确覆盖率
- 适合需要可靠置信区间的研究者使用
我们为随机森林及相关集成方法构建了经验似然(EL)框架,提供一种基于似然的统计不确定性量化方法。利用集成预测中固有的不完整U统计量结构,我们构造了一个渐近服从卡方分布的EL统计量,前提是子采样因不完整性导致的稀疏性不过分严重。在更稀疏的子采样情形下,原EL统计量会过度覆盖,因丧失枢轴性;为此我们提出一种简单调整的修正版EL,恢复其枢轴性。该方法保持了经验似然的核心性质,同时计算高效。针对诚实随机森林的理论分析与模拟实验表明,修正后的EL相比现有推断方法具有更准确的覆盖率和实际可靠性。
原文摘要 · Abstract (English)
We develop an empirical likelihood (EL) framework for random forests and related ensemble methods, providing a likelihood-based approach to quantify their statistical uncertainty. Exploiting the incomplete $U$-statistic structure inherent in ensemble predictions, we construct an EL statistic that is asymptotically chi-squared when subsampling induced by incompleteness is not overly sparse. Under sparser subsampling regimes, the EL statistic tends to over-cover due to loss of pivotality; we therefore propose a modified EL that restores pivotality through a simple adjustment. Our method retains key properties of EL while remaining computationally efficient. Theory for honest random forests and simulations demonstrate that modified EL achieves accurate coverage and practical reliability relative to existing inference methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。