对比多种模型的SHAP分析,揭示解释性差异。
A comparative analysis of machine learning models in SHAP analysis

- 在不同机器学习模型上系统分析SHAP值解释效果。
- 发现模型类型显著影响特征贡献度的解读方式。
- 提出多分类场景下的新型水桶图可视化方法。
在数据与技术飞速发展的时代,大型黑箱模型因其处理海量数据和学习复杂模式的能力而成为主流。然而,这些方法缺乏可解释性,难以阐明预测过程,使其在高风险场景中使用时存在信任问题。SHapley Additive exPlanations(SHAP)作为一种日益流行的可解释人工智能方法,能够基于原始特征解释模型预测。对每个样本和特征,SHAP值量化了该特征对预测结果的贡献。分析这些值可深入理解模型决策机制,为数据驱动的解决方案提供支持。但SHAP值的解释具有模型依赖性,尚无通用分析流程。为此,本文详细研究了不同机器学习模型和数据集上的SHAP分析。通过揭示其内在细节与细微差别,旨在帮助分析师在这一探索不足的领域中更有效地开展工作。同时,本文提出一种针对多分类问题的水桶图新泛化形式。
原文摘要 · Abstract (English)
In this growing age of data and technology, large black-box models are becoming the norm due to their ability to handle vast amounts of data and learn incredibly complex data patterns. The deficiency of these methods, however, is their inability to explain the prediction process, making them untrustworthy and their use precarious in high-stakes situations. SHapley Additive exPlanations (SHAP) analysis is an explainable AI method growing in popularity for its ability to explain model predictions in terms of the original features. For each sample and feature in the data set, an associated SHAP value quantifies the contribution of that feature to the prediction of that sample. Analysis of these SHAP values provides valuable insight into the model's decision-making process, which can be leveraged to create data-driven solutions. The interpretation of these SHAP values, however, is model-dependent, so there does not exist a universal analysis procedure. To aid in these efforts, we present a detailed investigation of SHAP analysis across various machine learning models and data sets. In uncovering the details and nuance behind SHAP analysis, we hope to empower analysts in this less-explored territory. We also present a novel generalization of the waterfall plot to the multi-classification problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。