arXiv:2601.04977cs.LGcs.AI2026-01被引 1

论文定义并检测了反事实解释中的‘挑拣’行为,发现难以被外部审计识别。

On the Definition and Detection of Cherry-Picking in Counterfactual Explanations

  • 基于生成过程与效用函数,形式化定义反事实解释的可接受空间。
  • 三种访问权限下检测能力均极有限,挑拣解释与正常解释难区分。
  • 建议优先保障流程标准化,而非事后检测,适合可信AI开发者参考。

反事实解释广泛用于说明模型预测改变所需输入的变动方式。对于单个实例,可能存在多种有效反事实解释,这为解释提供方选择性呈现有利行为、隐藏问题行为提供了可能。本文从生成过程定义的可接受解释空间和效用函数出发,形式化定义了反事实解释中的挑拣行为,并研究外部审计者在三种访问级别(完全流程访问、部分流程访问、仅解释访问)下的检测能力。结果表明,即使拥有完整流程访问,挑拣解释也难以与非挑拣解释区分,因为有效反事实的多样性及解释规范的灵活性提供了足够的自由度来掩盖刻意选择。实证显示,这种变异性通常远超挑拣行为对标准质量指标(如接近度、合理性、稀疏性)的影响,导致挑拣解释在统计上与基线解释无法区分。因此,建议优先通过可复现性、标准化和流程约束构建防护机制,而非依赖事后检测。本文为算法开发者、解释提供方和审计者提供具体建议。

原文摘要 · Abstract (English)

Counterfactual explanations are widely used to communicate how inputs must change for a model to alter its prediction. For a single instance, many valid counterfactuals can exist, which leaves open the possibility for an explanation provider to cherry-pick explanations that better suit a narrative of their choice, highlighting favourable behaviour and withholding examples that reveal problematic behaviour. We formally define cherry-picking for counterfactual explanations in terms of an admissible explanation space, specified by the generation procedure, and a utility function. We then study to what extent an external auditor can detect such manipulation. Considering three levels of access to the explanation process: full procedural access, partial procedural access, and explanation-only access, we show that detection is extremely limited in practice. Even with full procedural access, cherry-picked explanations can remain difficult to distinguish from non cherry-picked explanations, because the multiplicity of valid counterfactuals and flexibility in the explanation specification provide sufficient degrees of freedom to mask deliberate selection. Empirically, we demonstrate that this variability often exceeds the effect of cherry-picking on standard counterfactual quality metrics such as proximity, plausibility, and sparsity, making cherry-picked explanations statistically indistinguishable from baseline explanations. We argue that safeguards should therefore prioritise reproducibility, standardisation, and procedural constraints over post-hoc detection, and we provide recommendations for algorithm developers, explanation providers, and auditors.

反事实解释可信AI审计可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。