arXiv:2410.05419cs.LGcs.AI2024-10被引 4

用最优传输优化反事实解释,让修改更少却效果相同

Joint Distribution-Informed Shapley Values for Sparse Counterfactual Explanations

  • 通过最优传输构建事实与反事实的耦合关系
  • 在4个数据集上仅需原方案26%~45%的特征修改
  • 适合追求简洁可操作解释的开发者和研究者

反事实解释旨在揭示输入微小变动如何改变模型预测,但现有方法常过度修改特征,降低清晰度与可操作性。本文提出无需依赖模型或生成器的后处理框架COLA,通过最优传输(OT)建立事实与反事实集合间的耦合,并基于此计算一种改进的Shapley值(p-SHAP),以选取最小修改集并保持目标预测效果。理论上,该方法最小化事实与反事实结果间W1距离的上界,在弱条件下保证修正后的反事实不会比原始更远离事实。实验证明,在四个数据集、十二种模型及五种生成器上,COLA在保持相同预测效果的同时,仅需原方案26%~45%的特征修改;小规模基准测试显示其接近最优。

原文摘要 · Abstract (English)

Counterfactual explanations (CE) aim to reveal how small input changes flip a model's prediction, yet many methods modify more features than necessary, reducing clarity and actionability. We introduce \emph{COLA}, a model- and generator-agnostic post-hoc framework that refines any given CE by computing a coupling via optimal transport (OT) between factual and counterfactual sets and using it to drive a Shapley-based attribution (\emph{$p$-SHAP}) that selects a minimal set of edits while preserving the target effect. Theoretically, OT minimizes an upper bound on the $W_1$ divergence between factual and counterfactual outcomes and that, under mild conditions, refined counterfactuals are guaranteed not to move farther from the factuals than the originals. Empirically, across four datasets, twelve models, and five CE generators, COLA achieves the same target effects with only 26--45\% of the original feature edits. On a small-scale benchmark, COLA shows near-optimality.

反事实解释最优传输可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。