通过社区检测发现解释图中协同影响预测的特征模块
Community Detection on Model Explanation Graphs for Explainable AI
- 构建实例级归因的解释图,用社区检测找协同特征组
- 发现特征相关性,定位模型偏差与冗余模块,提升调试效率
- 适合关注模型可解释性与公平性的研究人员
特征归因方法(如SHAP、LIME)虽能解释单个预测,但常忽略高阶结构:共同作用的特征集合。我们提出影响模块(MoI)框架,(i) 从实例级归因构建模型解释图,(ii) 应用社区检测识别共同影响预测的特征模块,(iii) 量化这些模块与偏差、冗余、因果模式的关系。在合成与真实数据集上,MoI揭示了相关特征组,通过模块级消融改进模型调试,并将偏差暴露定位到特定模块。我们发布了稳定性与协同度指标、参考实现及评估协议,用于基准测试XAI中的模块发现。
原文摘要 · Abstract (English)
Feature-attribution methods (e.g., SHAP, LIME) explain individual predictions but often miss higher-order structure: sets of features that act in concert. We propose Modules of Influence (MoI), a framework that (i) constructs a model explanation graph from per-instance attributions, (ii) applies community detection to find feature modules that jointly affect predictions, and (iii) quantifies how these modules relate to bias, redundancy, and causality patterns. Across synthetic and real datasets, MoI uncovers correlated feature groups, improves model debugging via module-level ablations, and localizes bias exposure to specific modules. We release stability and synergy metrics, a reference implementation, and evaluation protocols to benchmark module discovery in XAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。