用高斯过程构建可解释的表格数据自编码器,生成更真实、多样化的反事实样本。
An Explainable Gaussian Process Auto-encoder for Tabular Data
- 基于高斯过程设计自编码器,参数少且不易过拟合
- 提出新密度估计器,能搜索分布内反事实样本
- 自动优化正则化率,适合高风险场景可解释性需求
可解释机器学习在高风险场景中备受关注,其中反事实解释成为解释黑箱模型的重要工具。近期进展利用生成模型(如自编码器)提升解释能力。本文提出一种基于高斯过程的新型自编码器架构,用于生成反事实样本。该模型参数更少,减少过拟合风险。同时引入新的密度估计器,支持在数据分布内搜索反事实样本,并提出算法自动选择最优正则化率。我们在多个大规模表格数据集上测试该方法,与现有自编码器基线对比,结果表明该方法能生成多样化且分布一致的反事实样本。
原文摘要 · Abstract (English)
Explainable machine learning has attracted much interest in the community where the stakes are high. Counterfactual explanations methods have become an important tool in explaining a black-box model. The recent advances have leveraged the power of generative models such as an autoencoder. In this paper, we propose a novel method using a Gaussian process to construct the auto-encoder architecture for generating counterfactual samples. The resulting model requires fewer learnable parameters and thus is less prone to overfitting. We also introduce a novel density estimator that allows for searching for in-distribution samples. Furthermore, we introduce an algorithm for selecting the optimal regularization rate on density estimator while searching for counterfactuals. We experiment with our method in several large-scale tabular datasets and compare with other auto-encoder-based methods. The results show that our method is capable of generating diversified and in-distribution counterfactual samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。