用反事实解释聚类结果,让模型决策更透明。
Counterfactual Explanations for k-means and Gaussian Clustering
- 定义可解释的反事实规则,约束合理性和可行性。
- 针对k均值和高斯聚类,分别给出解析解与单参数数值解。
- 适合需要理解聚类归属原因的研究者与开发者。
反事实解释已被证明是阐明分类器决策的有效方法,但在聚类领域尚未得到应用。本文提出将反事实用于解释基于模型的聚类结果。首先,我们给出了适用于模型聚类的反事实通用定义,包含合理性与可行性约束。接着,针对采用欧氏距离的k均值和高斯聚类,研究反事实生成问题。输入包括原始样本、目标簇、指示可修改或不可修改特征的二值掩码,以及指定反事实应位于簇边界多远的合理性因子。在k均值情况下,我们推导出最优解的解析公式;在高斯聚类情形(考虑完整、对角或球形协方差)下,方法需求解仅含一个参数的非线性方程。通过示例和定量实验验证了该方法的优势。
原文摘要 · Abstract (English)
Counterfactuals have been recognized as an effective approach to explain classifier decisions. Nevertheless, they have not yet been considered in the context of clustering. In this work, we propose the use of counterfactuals to explain clustering solutions. First, we present a general definition for counterfactuals for model-based clustering that includes plausibility and feasibility constraints. Then we consider the counterfactual generation problem for k-means and Gaussian clustering assuming Euclidean distance. Our approach takes as input the factual, the target cluster, a binary mask indicating actionable or immutable features and a plausibility factor specifying how far from the cluster boundary the counterfactual should be placed. In the k-means clustering case, analytical mathematical formulas are presented for computing the optimal solution, while in the Gaussian clustering case (assuming full, diagonal, or spherical covariances) our method requires the numerical solution of a nonlinear equation with a single parameter only. We demonstrate the advantages of our approach through illustrative examples and quantitative experimental comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。