arXiv:2507.23767stat.MLcs.LG2025-07

从少量统计量重建贝塔分布,提升随机森林分类效果。

Closed-Form Beta Distribution Estimation from Sparse Statistics with Random Forest Implicit Regularization

  • 用极值、均值、中位数匹配构造闭式贝塔分布估计器。
  • 恢复的分布参数使时间序列分类准确率提升,且与分布距离相关。
  • 零方差特征隐式正则化,生成更深更丰富的决策树,适合数据稀疏场景。

本文通过三项核心贡献推动稀疏数据下的分布重构与集成分类。首先,提出一种闭式估计器,仅凭最小值、最大值、均值和中位数即可重构缩放后的贝塔分布,通过复合分位数与矩匹配实现;将恢复的参数(α, β)作为随机森林特征,显著提升时间序列快照的成对分类性能,验证了分布重建的保真度。其次,建立了分类准确率与分布接近度之间的联系,推导出总变差距离与Jensen-Shannon散度的误差界,后者呈现二次收敛特性。第三,发现零方差特征可作为隐式正则项,提高中等排名预测因子的选择概率,促使树结构更深、更具多样性。以SeatGeek票价数据集为主要应用,展示分布重建与事件级分类能力,并揭示二级票务市场的结构动态;UCI手写数字数据集进一步证实该正则化效应的普适性。整体上,研究为从稀疏分布快照到闭式估计与集成模型精度提升提供了实用路径,且可靠性通过隐式正则化得到增强。

原文摘要 · Abstract (English)

This work advances distribution recovery from sparse data and ensemble classification through three main contributions. First, we introduce a closed-form estimator that reconstructs scaled beta distributions from limited statistics (minimum, maximum, mean, and median) via composite quantile and moment matching. The recovered parameters $(α,β)$, when used as features in Random Forest classifiers, improve pairwise classification on time-series snapshots, validating the fidelity of the recovered distributions. Second, we establish a link between classification accuracy and distributional closeness by deriving error bounds that constrain total variation distance and Jensen-Shannon divergence, the latter exhibiting quadratic convergence. Third, we show that zero-variance features act as an implicit regularizer, increasing selection probability for mid-ranked predictors and producing deeper, more varied trees. A SeatGeek pricing dataset serves as the primary application, illustrating distributional recovery and event-level classification while situating these methods within the structure and dynamics of the secondary ticket marketplace. The UCI handwritten digits dataset confirms the broader regularization effect. Overall, the study outlines a practical route from sparse distributional snapshots to closed-form estimation and improved ensemble accuracy, with reliability enhanced through implicit regularization.

分布估计随机森林隐式正则化稀疏数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。