用SHAP值聚类分析样本预测原因,揭示相同结果的不同成因路径。
SHAP-Based Supervised Clustering for Sample Classification and the Generalized Waterfall Plot
- 基于SHAP值聚类,按预测原因分组样本
- 在阿尔茨海默病数据上验证方法有效性
- 提出多分类用水桶图的通用扩展
在数据与技术迅猛发展的时代,大型黑箱模型因其处理海量数据和学习复杂输入输出关系的能力已成为主流。然而,这些方法缺乏可解释性,难以说明预测过程,使其在高风险场景中不可靠。SHapley Additive exPlanations(SHAP)是一种日益流行的可解释人工智能方法,能以原始特征为单位解释模型预测。对每个样本和特征,计算其对应的SHAP值,量化该特征对预测的贡献。对这些SHAP值进行聚类,可揭示数据内在结构:将不仅预测相同,且因相似原因而预测相同的样本归为一类。这有助于映射不同样本达成相同预测的多种路径。本文通过模拟实验和基于阿尔茨海默病神经影像计划(ADNI)数据库的病例研究展示该方法。同时提出一种适用于多分类任务的水桶图通用化形式。
原文摘要 · Abstract (English)
In this growing age of data and technology, large black-box models are becoming the norm due to their ability to handle vast amounts of data and learn incredibly complex input-output relationships. The deficiency of these methods, however, is their inability to explain the prediction process, making them untrustworthy and their use precarious in high-stakes situations. SHapley Additive exPlanations (SHAP) analysis is an explainable AI method growing in popularity for its ability to explain model predictions in terms of the original features. For each sample and feature in the data set, we associate a SHAP value that quantifies the contribution of that feature to the prediction of that sample. Clustering these SHAP values can provide insight into the data by grouping samples that not only received the same prediction, but received the same prediction for similar reasons. In doing so, we map the various pathways through which distinct samples arrive at the same prediction. To showcase this methodology, we present a simulated experiment in addition to a case study in Alzheimer's disease using data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. We also present a novel generalization of the waterfall plot for multi-classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。