arXiv:2606.04209cs.LG2026-06

研究模型决策边界与数据分布的关系,揭示解释性差异的根源

A Geometric View of Counterfactual Behavior: Interaction of Boundary Proximity and Local Support

论文配图:A Geometric View of Counterfactual Behavior: Interaction of Boundary Proximity and Local Support
图 1 · 摘自论文原文
  • 用局部搜索探针分析预训练模型的反事实行为
  • 相同准确率下,不同分类头导致反事实可实现性差异显著
  • 边界靠近数据且支持区域充足时,反事实更可行

反事实解释通过微小且语义合理的输入变化来改变模型预测,广泛用于机器学习系统的可解释性与审计。在现代视觉、语言及多模态系统中,预训练编码器将输入映射到表示空间,下游分类头在该空间内设定决策边界。因此,附近反事实的可行性与距离取决于边界相对于数据的位置。然而,具有相似预测性能的模型在反事实可实现性上可能有显著差异。本研究通过标准化局部搜索探针,在多个预训练编码器和线性分类头上进行测试。结果表明,尽管预测性能相似,模型的反事实行为仍存在显著差异;在固定表示下,仅改变分类头即可影响反事实结果,而预测性能基本不变。这种差异由决策边界邻近性与局部数据支持的相互作用决定,共同决定了预测变化是否既可行又位于数据支持区域内,并可提升固定模型内的反事实搜索效率。这些发现表明,反事实行为是独立于预测性能的一个维度,且可在不改变准确率的情况下被调整,对模型选择、鲁棒性及反事实方法可靠性具有重要启示。

原文摘要 · Abstract (English)

Counterfactual explanations seek small, semantically meaningful changes to an input that alter a model's prediction, and are widely used to interpret and audit machine learning systems. In modern vision, language, and multimodal systems, pretrained encoders map inputs to representation spaces, and downstream classifier heads impose decision boundaries within those spaces. As a result, the feasibility and distance of nearby counterfactuals depend on boundary placement relative to the data. Yet models with similar predictive performance can differ substantially in whether such changes are achievable and how far representations must move. This work examines this variation using a standardized local search probe across several pretrained encoders and linear classifier heads. Results show that despite similar predictive performance, models differ substantially in their counterfactual behavior. Under fixed representations, varying only the classifier head alters counterfactual outcomes while leaving predictive performance largely unchanged. This variation is explained by the interaction of decision-boundary proximity and local data support, which jointly determine whether prediction changes are both feasible and lie in regions supported by the data, and can also improve counterfactual search within fixed models. Together, these findings identify counterfactual behavior as a distinct dimension beyond predictive performance and show that it can be altered without changing accuracy, with implications for model selection, robustness, and the reliability of counterfactual methods.

反事实解释可解释性决策边界模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。