arXiv:2502.11509cs.LGcs.AI2025-02被引 3

用扩散自编码器和聚类生成多样反事实解释

DifCluE: Generating Counterfactual Explanations with Diffusion Autoencoders and modal clustering

  • 在隐空间聚类发现类别内不同模式方向
  • 生成多个差异明显且有意义的反事实样本
  • 适合需要可解释性的AI决策场景

为同一类别中的不同数据模式生成多个反事实解释是一项重大挑战,因为这些模式虽各异却共享相同分类结果。扩散概率模型(DPMs)展现出捕捉数据分布底层模式的强大能力。本文提出DifCluE方法,利用扩散自编码器在隐空间进行聚类,识别出类别内部不同模式对应的生成方向,从而生成多样化且语义合理的反事实解释。实验表明,DifCluE在生成多反事实解释方面优于当前最先进方法,显著提升了模型可解释性。

原文摘要 · Abstract (English)

Generating multiple counterfactual explanations for different modes within a class presents a significant challenge, as these modes are distinct yet converge under the same classification. Diffusion probabilistic models (DPMs) have demonstrated a strong ability to capture the underlying modes of data distributions. In this paper, we harness the power of a Diffusion Autoencoder to generate multiple distinct counterfactual explanations. By clustering in the latent space, we uncover the directions corresponding to the different modes within a class, enabling the generation of diverse and meaningful counterfactuals. We introduce a novel methodology, DifCluE, which consistently identifies these modes and produces more reliable counterfactual explanations. Our experimental results demonstrate that DifCluE outperforms the current state-of-the-art in generating multiple counterfactual explanations, offering a significant advancement in model interpretability.

反事实解释扩散模型可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。