让模型在训练中学会生成合理且可操作的反事实解释
Counterfactual Training: Teaching Models Plausible and Actionable Explanations
- 训练时引入反事实解释,让模型直接学习可解释性
- 生成的反事实既合理又可操作,且提升模型抗干扰能力
- 适合需要高可信度解释的医疗、金融等决策系统
我们提出一种新型训练方法——反事实训练,利用反事实解释增强模型的可解释能力。反事实解释能说明输入如何改变才能使模型输出期望结果,但现有研究多聚焦于事后生成满足数据合理性与特征可变性约束的反事实。本文则将反事实纳入训练过程,使模型在学习过程中最小化其表征与合理、可操作解释之间的差异。实证与理论分析表明,该方法能训练出天然具备优质反事实解释能力的模型,并同时提升对抗鲁棒性。
原文摘要 · Abstract (English)
We propose a novel training regime termed counterfactual training that leverages counterfactual explanations to increase the explanatory capacity of models. Counterfactual explanations have emerged as a popular post-hoc explanation method for opaque machine learning models: they inform how factual inputs would need to change in order for a model to produce some desired output. To be useful in real-world decision-making systems, counterfactuals should be plausible with respect to the underlying data and actionable with respect to the feature mutability constraints. Much existing research has therefore focused on developing post-hoc methods to generate counterfactuals that meet these desiderata. In this work, we instead hold models directly accountable for the desired end goal: counterfactual training employs counterfactuals during the training phase to minimize the divergence between learned representations and plausible, actionable explanations. We demonstrate empirically and theoretically that our proposed method facilitates training models that deliver inherently desirable counterfactual explanations and additionally exhibit improved adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。