用黎曼几何优化生成反事实样本,让解释更自然可信。
Counterfactual Explanations via Riemannian Latent Space Traversal
- 通过解码器和分类器拉回黎曼度量,捕捉数据复杂结构
- 在真实表格数据上生成高保真、自然的反事实路径
- 适合需要可行动解释的模型可解释性研究者
深度模型日益复杂,亟需理解其预测机制。反事实解释为从业者提供可操作的洞察。以往方法依赖生成模型的潜在空间遍历,但这些潜在空间通常过于简化,数据分布的复杂性主要存在于解码器而非潜在嵌入中。若不考虑非线性解码器而直接遍历潜在空间,会导致反事实轨迹不自然。本文提出基于解码器与待检分类器拉回的黎曼度量来生成反事实解释。该度量编码了数据与学习表征的复杂几何结构,使反事实轨迹具备鲁棒性和高保真度。实验在真实世界表格数据集上验证了该方法的有效性。
原文摘要 · Abstract (English)
The adoption of increasingly complex deep models has fueled an urgent need for insight into how these models make predictions. Counterfactual explanations form a powerful tool for providing actionable explanations to practitioners. Previously, counterfactual explanation methods have been designed by traversing the latent space of generative models. Yet, these latent spaces are usually greatly simplified, with most of the data distribution complexity contained in the decoder rather than the latent embedding. Thus, traversing the latent space naively without taking the nonlinear decoder into account can lead to unnatural counterfactual trajectories. We introduce counterfactual explanations obtained using a Riemannian metric pulled back via the decoder and the classifier under scrutiny. This metric encodes information about the complex geometric structure of the data and the learned representation, enabling us to obtain robust counterfactual trajectories with high fidelity, as demonstrated by our experiments in real-world tabular datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。