arXiv:2505.23225cs.LG2025-05

模型过拟合越严重,越容易生成反事实解释,导致泛化能力下降。

Generalizability vs. Counterfactual Explainability Trade-Off

  • 提出ε-有效反事实概率,衡量数据点附近扰动引发标签变化的概率
  • 理论证明过拟合时反事实概率上升,揭示泛化与可解释性间的权衡
  • 该指标可作过拟合的量化代理,适合关注模型鲁棒性的研究者

本文研究监督学习中模型泛化能力与反事实可解释性之间的关系。引入ε-有效反事实概率(ε-VCP)——即在数据点ε邻域内找到使标签改变的扰动的概率。理论上分析了ε-VCP与模型决策边界几何结构的关系,发现ε-VCP随模型过拟合程度增加而上升。研究结果建立了泛化性能差与反事实生成容易之间的严格联系,揭示了泛化与反事实可解释性之间的内在权衡。实验结果验证了理论,表明ε-VCP可作为量化刻画过拟合的有效代理指标。

原文摘要 · Abstract (English)

In this work, we investigate the relationship between model generalization and counterfactual explainability in supervised learning. We introduce the notion of $\varepsilon$-valid counterfactual probability ($\varepsilon$-VCP) -- the probability of finding perturbations of a data point within its $\varepsilon$-neighborhood that result in a label change. We provide a theoretical analysis of $\varepsilon$-VCP in relation to the geometry of the model's decision boundary, showing that $\varepsilon$-VCP tends to increase with model overfitting. Our findings establish a rigorous connection between poor generalization and the ease of counterfactual generation, revealing an inherent trade-off between generalization and counterfactual explainability. Empirical results validate our theory, suggesting $\varepsilon$-VCP as a practical proxy for quantitatively characterizing overfitting.

模型可解释性过拟合反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。