arXiv:2502.07153cs.LGcs.AI2025-02被引 1

数据特性决定解释效果,选对方法才能可靠分析决策树模型。

Feature Importance Depends on Properties of the Data: Towards Choosing the Correct Explanations for Your Data and Decision Trees based Models

  • 用合成数据测试不同解释方法在决策树上的表现。
  • 发现特征重要性数值和符号受数据特性显著影响。
  • 提醒用户根据数据特点选择合适的解释方法。

为确保机器学习模型解释的可靠性,必须明确各类解释方法的优势与局限及其适用场景。然而,当前对解释方法何时何地有效的理解仍不充分。为此,我们通过构建具有特定属性的合成数据集,开展全面的实证评估。研究目标是衡量局部解释方法对决策树模型预测结果所提供的特征重要性估计的质量。通过对合成数据集及公开二分类数据集的分析,我们发现不同解释方法产生的特征重要性估计在数值大小和符号上存在显著差异,且这些估计对数据中的特定属性敏感。尽管部分模型超参数对特征重要性分配影响不大,但每种解释方法在特定情境下均有其局限性。本研究揭示了这些局限,并为不同场景下解释方法的选择与可靠性提供了关键洞见。

原文摘要 · Abstract (English)

In order to ensure the reliability of the explanations of machine learning models, it is crucial to establish their advantages and limits and in which case each of these methods outperform. However, the current understanding of when and how each method of explanation can be used is insufficient. To fill this gap, we perform a comprehensive empirical evaluation by synthesizing multiple datasets with the desired properties. Our main objective is to assess the quality of feature importance estimates provided by local explanation methods, which are used to explain predictions made by decision tree-based models. By analyzing the results obtained from synthetic datasets as well as publicly available binary classification datasets, we observe notable disparities in the magnitude and sign of the feature importance estimates generated by these methods. Moreover, we find that these estimates are sensitive to specific properties present in the data. Although some model hyper-parameters do not significantly influence feature importance assignment, it is important to recognize that each method of explanation has limitations in specific contexts. Our assessment highlights these limitations and provides valuable insight into the suitability and reliability of different explanatory methods in various scenarios.

模型解释决策树特征重要性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。