arXiv:2501.16756cs.LG2025-01被引 16

优化随机森林可媲美先进校准方法,无需大量数据

Random Forest Calibration

  • 系统测试多种校准方法,发现调优后的随机森林表现优异
  • 在有限数据下,优化的随机森林校准效果优于传统方法
  • 适合数据量小、需可靠概率估计的场景

随机森林(RF)分类器常被认为相比其他机器学习方法具有较好的概率校准能力。现有研究指出,传统校准方法如等熵回归在缺乏大量校准数据时难以显著提升RF的概率估计,这在数据有限的情况下成为障碍。然而,尚无全面研究验证该说法并系统比较适用于RF的先进校准方法。为此,我们考察了涵盖缩放技术到高级算法在内的广泛校准方法,基于合成与真实数据集展开分析,揭示了RF概率估计的复杂性,评估了超参数影响,并系统比较了各类方法。结果表明,经过良好优化的随机森林在性能上可达到甚至超过主流校准方法。

原文摘要 · Abstract (English)

The Random Forest (RF) classifier is often claimed to be relatively well calibrated when compared with other machine learning methods. Moreover, the existing literature suggests that traditional calibration methods, such as isotonic regression, do not substantially enhance the calibration of RF probability estimates unless supplied with extensive calibration data sets, which can represent a significant obstacle in cases of limited data availability. Nevertheless, there seems to be no comprehensive study validating such claims and systematically comparing state-of-the-art calibration methods specifically for RF. To close this gap, we investigate a broad spectrum of calibration methods tailored to or at least applicable to RF, ranging from scaling techniques to more advanced algorithms. Our results based on synthetic as well as real-world data unravel the intricacies of RF probability estimates, scrutinize the impacts of hyper-parameters, compare calibration methods in a systematic way. We show that a well-optimized RF performs as well as or better than leading calibration approaches.

随机森林概率校准机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。