arXiv:2409.06439cs.LGstat.CO2024-09被引 2

将可解释随机森林方法拓展至回归任务,提升模型透明度。

Extending Explainable Ensemble Trees (E2Tree) to regression contexts

  • 基于相似性度量,可视化预测变量与响应变量关系
  • 首次实现对回归型随机森林的完整可解释性分析
  • 适合需要透明决策过程的研究者和从业者

集成学习方法如随机森林通过聚合多个弱学习器显著提升了监督学习的预测精度。然而,这些方法往往缺乏透明性,阻碍用户理解模型如何做出预测。可解释集成树(E2Tree)是一种新型随机森林解释方法,能以图形方式展示响应变量与预测变量之间的关系。其显著特点在于不仅考虑预测变量对响应变量的影响,还通过计算和使用相似性度量来捕捉预测变量间的关联。该方法最初仅适用于分类任务。本文将其扩展至回归场景。为验证算法的解释能力,我们在真实数据集上展示了其应用效果。

原文摘要 · Abstract (English)

Ensemble methods such as random forests have transformed the landscape of supervised learning, offering highly accurate prediction through the aggregation of multiple weak learners. However, despite their effectiveness, these methods often lack transparency, impeding users' comprehension of how RF models arrive at their predictions. Explainable ensemble trees (E2Tree) is a novel methodology for explaining random forests, that provides a graphical representation of the relationship between response variables and predictors. A striking characteristic of E2Tree is that it not only accounts for the effects of predictor variables on the response but also accounts for associations between the predictor variables through the computation and use of dissimilarity measures. The E2Tree methodology was initially proposed for use in classification tasks. In this paper, we extend the methodology to encompass regression contexts. To demonstrate the explanatory power of the proposed algorithm, we illustrate its use on real-world datasets.

可解释性随机森林回归分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。