arXiv:2505.16583cs.LG2025-05

用合理反事实数据训练,能有效消除模型对虚假关联的依赖。

Training on Plausible Counterfactuals Removes Spurious Correlations

  • 通过生成合理反事实样本并标记为错误标签来训练模型
  • 模型在真实数据上准确率高,且虚假相关性显著降低
  • 适合关注模型公平性与鲁棒性的研究者

合理反事实解释(p-CFEs)是轻微扰动输入以改变分类结果,同时仍符合数据分布的扰动。本研究证明,可使用由p-CFEs诱导的错误标签训练分类器,使其对原始未扰动输入做出正确预测。此前研究仅在对抗扰动下验证该方法,本文将其拓展至p-CFEs。实验表明,基于p-CFEs的训练更有效:所得分类器不仅在分布内准确率高,且对虚假相关性的偏差显著减小。

原文摘要 · Abstract (English)

Plausible counterfactual explanations (p-CFEs) are perturbations that minimally modify inputs to change classifier decisions while remaining plausible under the data distribution. In this study, we demonstrate that classifiers can be trained on p-CFEs labeled with induced \emph{incorrect} target classes to classify unperturbed inputs with the original labels. While previous studies have shown that such learning is possible with adversarial perturbations, we extend this paradigm to p-CFEs. Interestingly, our experiments reveal that learning from p-CFEs is even more effective: the resulting classifiers achieve not only high in-distribution accuracy but also exhibit significantly reduced bias with respect to spurious correlations.

反事实学习虚假相关模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。