arXiv:2512.10878cs.LG2025-12中稿 · ICML

用反事实样本优化模型重建,提升逼近效果

Classifier Reconstruction Through Counterfactual-Aware Wasserstein Prototypes

  • 融合原始数据与反事实样本来构建类原型
  • 在少样本下提升代理模型与目标模型的匹配度
  • 适合需要模型可解释性重建的场景

反事实解释通过识别最小输入变化来提供可操作的洞察,不仅能增强模型可解释性,还可用于模型重建——训练一个代理模型以复现目标模型的行为。本文发现,反事实样本虽靠近决策边界,但对两类都具有信息价值,尽管代表性较弱。在标注数据有限的场景下,这一特性尤为关键。我们提出一种方法:将原始样本与反事实样本结合,利用Wasserstein均值逼近每个类的原型,从而保留类间分布结构。该方法提升了代理模型质量,并缓解了将反事实样本简单视为普通训练数据时常见的决策边界偏移问题。多个数据集上的实证结果表明,该方法显著提高了代理模型与目标模型之间的保真度。

原文摘要 · Abstract (English)

Counterfactual explanations provide actionable insights by identifying minimal input changes required to achieve a desired model prediction. Beyond their interpretability benefits, counterfactuals can also be leveraged for model reconstruction, where a surrogate model is trained to replicate the behavior of a target model. In this work, we demonstrate that model reconstruction can be significantly improved by recognizing that counterfactuals, which typically lie close to the decision boundary, can serve as informative though less representative samples for both classes. This is particularly beneficial in settings with limited access to labeled data. We propose a method that integrates original data samples with counterfactuals to approximate class prototypes using the Wasserstein barycenter, thereby preserving the underlying distributional structure of each class. This approach enhances the quality of the surrogate model and mitigates the issue of decision boundary shift, which commonly arises when counterfactuals are naively treated as ordinary training instances. Empirical results across multiple datasets show that our method improves fidelity between the surrogate and target models, validating its effectiveness.

模型重建反事实解释分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。