用合理反事实数据训练,能有效消除模型对虚假关联的依赖。
Training on Plausible Counterfactuals Removes Spurious Correlations
- 通过生成合理反事实样本并标记为错误标签来训练模型
- 模型在真实数据上准确率高,且虚假相关性显著降低
- 适合关注模型公平性与鲁棒性的研究者
合理反事实解释(p-CFEs)是轻微扰动输入以改变分类结果,同时仍符合数据分布的扰动。本研究证明,可使用由p-CFEs诱导的错误标签训练分类器,使其对原始未扰动输入做出正确预测。此前研究仅在对抗扰动下验证该方法,本文将其拓展至p-CFEs。实验表明,基于p-CFEs的训练更有效:所得分类器不仅在分布内准确率高,且对虚假相关性的偏差显著减小。
原文摘要 · Abstract (English)
Plausible counterfactual explanations (p-CFEs) are perturbations that minimally modify inputs to change classifier decisions while remaining plausible under the data distribution. In this study, we demonstrate that classifiers can be trained on p-CFEs labeled with induced \emph{incorrect} target classes to classify unperturbed inputs with the original labels. While previous studies have shown that such learning is possible with adversarial perturbations, we extend this paradigm to p-CFEs. Interestingly, our experiments reveal that learning from p-CFEs is even more effective: the resulting classifiers achieve not only high in-distribution accuracy but also exhibit significantly reduced bias with respect to spurious correlations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。