arXiv:2410.19098stat.MLcs.LG2024-10被引 4

用浅层树构建可解释的集成模型,兼顾性能与透明度。

Inherently Interpretable Tree Ensemble Learning

  • 用浅层决策树做基学习器,使集成模型天然可解释
  • 实验表明新方法在解释性与预测性能间平衡更优
  • 适合需要透明决策的医疗、金融等场景

随机森林和梯度提升机等树集成模型因优异的预测性能被广泛应用。然而,由大量决策树构成的高性能集成模型缺乏足够的透明度与可解释性。本文证明:当使用浅层决策树作为基学习器时,集成学习算法不仅能通过广义加性模型的等价表示实现内在可解释性,有时还能获得更好的泛化性能。首先,提出一种解释算法,将树集成转换为具有内在可解释性的函数ANOVA表示;其次,提出两种增强可解释性的策略:训练阶段添加约束,以及事后效应剪枝。在模拟数据和真实数据集上的实验表明,相比基准方法,本方法在模型可解释性与预测性能之间取得了更优的权衡。

原文摘要 · Abstract (English)

Tree ensemble models like random forests and gradient boosting machines are widely used in machine learning due to their excellent predictive performance. However, a high-performance ensemble consisting of a large number of decision trees lacks sufficient transparency and explainability. In this paper, we demonstrate that when shallow decision trees are used as base learners, the ensemble learning algorithms can not only become inherently interpretable subject to an equivalent representation as the generalized additive models but also sometimes lead to better generalization performance. First, an interpretation algorithm is developed that converts the tree ensemble into the functional ANOVA representation with inherent interpretability. Second, two strategies are proposed to further enhance the model interpretability, i.e., by adding constraints in the model training stage and post-hoc effect pruning. Experiments on simulations and real-world datasets show that our proposed methods offer a better trade-off between model interpretation and predictive performance, compared with its counterpart benchmarks.

可解释性树模型集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。