arXiv:2512.11081stat.MLcs.LG2025-12

提出可证明的局部特征重要性方法,精准识别随机森林中影响单个预测的关键特征与交互。

Provable Recovery of Locally Important Signed Features and Interactions from Random Forest

  • 基于决策路径中的特征共现模式,融合全局与局部信息计算重要性
  • 在局部稀疏模型下,能一致恢复真实的重要特征和交互作用
  • 适用于个性化医疗等需个体化解释的场景,结果可信且可追溯

特征与交互重要性(FII)方法在监督学习中至关重要,用于评估输入变量及其交互对复杂预测模型的影响。在个性化医疗等领域,常需针对单个预测提供局部解释,而非全局重要性汇总。随机森林(RF)广泛应用于此类场景,现有可解释性方法通常利用树结构与分裂统计量提供模型特定洞察。然而,对随机森林局部FII方法的理论理解仍不充分,难以解释单个预测中高重要性分数的来源。本文提出一种新颖的局部、模型特定的FII方法,通过识别决策路径上特征的频繁共现,结合全局模式与特定测试点路径上的局部模式。我们证明,在局部尖峰稀疏(LSS)模型下,该方法能一致恢复真实的局部信号特征及其交互,并可判断是大值或小值驱动预测。通过模拟研究和真实数据案例验证了方法的有效性与理论结果的实用性。

原文摘要 · Abstract (English)

Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models. In many domains, such as personalized medicine, local interpretations for individual predictions are often required, rather than global scores summarizing overall feature importance. Random Forests (RFs) are widely used in these settings, and existing interpretability methods typically exploit tree structures and split statistics to provide model-specific insights. However, theoretical understanding of local FII methods for RF remains limited, making it unclear how to interpret high importance scores for individual predictions. We propose a novel, local, model-specific FII method that identifies frequent co-occurrences of features along decision paths, combining global patterns with those observed on paths specific to a given test point. We prove that our method consistently recovers the true local signal features and their interactions under a Locally Spike Sparse (LSS) model and also identifies whether large or small feature values drive a prediction. We illustrate the usefulness of our method and theoretical results through simulation studies and a real-world data example.

随机森林局部解释特征重要性可证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。