用沙普利值优化让视觉模型决策过程更透明,解释结果与模型行为一致。
Enhancing Interpretability for Vision Models via Shapley Value Optimization
- 训练时加入沙普利值估计作为辅助任务,实现图像块的公平贡献分配
- 在多个基准上达到当前最佳可解释性,且性能和兼容性几乎不受影响
- 适合需要高可信度解释的医疗、自动驾驶等安全敏感领域
深度神经网络在多个领域表现卓越,但其决策过程仍不透明。现有解释方法存在局限:事后解释方法难以真实反映模型行为,自解释模型则因特殊架构设计牺牲性能与兼容性。为此,我们提出一种新型自解释框架,在训练中引入沙普利值估计作为辅助任务,实现两个关键进步:1)将模型预测得分公平分配给图像块,确保解释与模型决策逻辑一致;2)通过微小结构修改提升可解释性,同时保持模型性能与兼容性。在多个基准上的大量实验表明,该方法在可解释性方面达到当前最优水平。
原文摘要 · Abstract (English)
Deep neural networks have demonstrated remarkable performance across various domains, yet their decision-making processes remain opaque. Although many explanation methods are dedicated to bringing the obscurity of DNNs to light, they exhibit significant limitations: post-hoc explanation methods often struggle to faithfully reflect model behaviors, while self-explaining neural networks sacrifice performance and compatibility due to their specialized architectural designs. To address these challenges, we propose a novel self-explaining framework that integrates Shapley value estimation as an auxiliary task during training, which achieves two key advancements: 1) a fair allocation of the model prediction scores to image patches, ensuring explanations inherently align with the model's decision logic, and 2) enhanced interpretability with minor structural modifications, preserving model performance and compatibility. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。