arXiv:2604.17494cs.LGcs.AI2026-04

用概率共识生成稳定可信的反事实解释,无需重训练。

A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations

论文配图:A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations
图 1 · 摘自论文原文
  • 基于模型集成的概率共识,建模数据分布与决策空间。
  • 仅需调整一个参数即可控制鲁棒性,且保持高可解释性。
  • 适合需要稳定解释的黑箱模型应用,如医疗、金融决策。

反事实解释(CFEs)对理解黑箱模型至关重要,但模型稍作改变后常失效。现有方法多局限于特定模型、调参成本高或鲁棒性控制僵化。本文提出新方法,联合建模数据分布与合理模型决策空间,以确保对模型变化的鲁棒性。通过模型集成的概率共识,训练一个条件归一化流,捕捉在不同分类器一致程度下的数据密度。推理时仅需一个可解释参数控制鲁棒性水平,即指定最小模型共识比例,无需重新训练生成模型。该方法有效将反事实解释推向既合理又稳定的区域。实验表明,本方法在实证鲁棒性上优于现有方法,同时在其他评估指标上表现良好。

原文摘要 · Abstract (English)

Counterfactual explanations (CFEs) are essential for interpreting black-box models, yet they often become invalid when models are slightly changed. Existing methods for generating robust CFEs are often limited to specific types of models, require costly tuning, or inflexible robustness controls. We propose a novel approach that jointly models the data distribution and the space of plausible model decisions to ensure robustness to model changes. Using a probabilistic consensus over a model ensemble, we train a conditional normalizing flow that captures the data density under varying levels of classifier agreement. At inference time, a single interpretable parameter controls the robustness level; it specifies the minimum fraction of models that should agree on the target class without retraining the generative model. Our method effectively pushes CFEs toward regions that are both plausible and stable across model changes. Experimental results demonstrate that our approach achieves superior empirical robustness while also maintaining good performance across other evaluation measures.

反事实解释鲁棒性生成模型模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。