提出决策树的观测多重性分解方法,揭示结构不稳是主要变异来源。
Decomposing Observational Multiplicity in Decision Trees: Leaf and Structural Regret
- 将决策树的预测多样性分解为叶节点和结构两部分误差。
- 实验证明结构误差贡献超过叶误差15倍以上,主导预测差异。
- 用误差指标做选择性预测,可将召回率从92%提升至100%。
许多机器学习任务存在多个表现相近的模型,称为预测多重性。其根源之一是观测多重性,源于标签收集的随机性:训练标签仅为真实概率的单次实现。尽管逻辑回归已有理论框架,但对非光滑、分段式模型如决策树的研究仍不足。本文提出决策树分类器的两种互补观测多重性概念:叶节点后悔(leaf regret)衡量固定叶内因样本有限带来的预测波动;结构后悔(structural regret)反映树结构本身不稳定的引发的变异性。我们正式分解了观测多重性,并提供统计保证。在多个信用风险评分数据集上的实验表明,理论分解与实际方差高度一致。值得注意的是,结构后悔是主要驱动因素,在某些数据集中其贡献超过叶后悔的15倍以上。此外,利用这些后悔指标作为选择性预测中的拒识机制,可有效识别任意区域,显著提升模型安全性,在最稳定子群体中将召回率从92%提升至100%。该研究建立了一个严谨的观测多重性量化框架,契合当前算法安全与可解释性进展。
原文摘要 · Abstract (English)
Many machine learning tasks admit multiple models that perform almost equally well, a phenomenon known as predictive multiplicity. A fundamental source of this multiplicity is observational multiplicity, which arises from the stochastic nature of label collection: observed training labels represent only a single realization of the underlying ground-truth probabilities. While theoretical frameworks for observational multiplicity have been established for logistic regression, their implications for non-smooth, partition-based models like decision trees remain underexplored. In this paper, we introduce two complementary notions of observational multiplicity for decision tree classifiers: leaf regret and structural regret. Leaf regret quantifies the intrinsic variability of predictions within a fixed leaf due to finite-sample noise, while structural regret captures variability induced by the instability of the learned tree structure itself. We provide a formal decomposition of observational multiplicity into these two components and establish statistical guarantees. Our experimental evaluation across diverse credit risk scoring datasets confirms the near-perfect alignment between our theoretical decomposition and the empirically observed variance. Notably, we find that structural regret is the primary driver of observational multiplicity, accounting for over 15 times the variability of leaf regret in some datasets. Furthermore, we demonstrate that utilizing these regret measures as an abstention mechanism in selective prediction can effectively identify arbitrary regions and improve model safety, elevating recall from 92% to 100% on the most stable sub-populations. These results establish a rigorous framework for quantifying observational multiplicity, aligning with recent advances in algorithmic safety and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。