arXiv:2503.05424cs.LGcs.CV2025-03中稿 · WACV-2026, 45 page…被引 2

通过渐进式干预揭示模型预测的因果驱动因素。

Locally Explaining Prediction Behavior via Gradual Interventions and Measuring Property Gradients

  • 利用图像编辑模型对语义属性进行渐进干预
  • 提出期望属性梯度幅值衡量预测影响程度
  • 适用于医疗诊断、训练分析等真实场景

深度学习模型虽具高预测性能,但缺乏内在可解释性,阻碍了对其预测行为的理解。现有局部可解释方法多关注关联关系,忽略因果驱动因素;而采用因果视角的方法主要提供全局模型级解释,无法确定这些因素是否适用于特定输入。为此,我们引入一种基于图像到图像编辑模型新进展的局部干预解释框架。该方法通过渐进干预语义属性,利用新提出的‘期望属性梯度幅值’得分量化其对模型预测的影响。我们在多种架构和任务上进行了广泛实证评估:首先在合成场景中验证其识别局部偏见的能力;随后应用于医学皮肤病变分类器,分析网络训练动态,并研究预训练CLIP模型在真实干预数据下的表现。结果表明,属性层面的干预解释能够揭示深度模型行为的新洞见。

原文摘要 · Abstract (English)

Deep learning models achieve high predictive performance but lack intrinsic interpretability, hindering our understanding of the learned prediction behavior. Existing local explainability methods focus on associations, neglecting the causal drivers of model predictions. Other approaches adopt a causal perspective but primarily provide global, model-level explanations. However, for specific inputs, it's unclear whether globally identified factors apply locally. To address this limitation, we introduce a novel framework for local interventional explanations by leveraging recent advances in image-to-image editing models. Our approach performs gradual interventions on semantic properties to quantify the corresponding impact on a model's predictions using a novel score, the expected property gradient magnitude. We demonstrate the effectiveness of our approach through an extensive empirical evaluation on a wide range of architectures and tasks. First, we validate it in a synthetic scenario and demonstrate its ability to locally identify biases. Afterward, we apply our approach to investigate medical skin lesion classifiers, analyze network training dynamics, and study a pre-trained CLIP model with real-life interventional data. Our results highlight the potential of interventional explanations on the property level to reveal new insights into the behavior of deep models.

可解释性因果推理图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。