arXiv:2512.10659cs.LG2025-12KDD被引 1

为局部离群点检测生成可解释的反事实样本

DCFO: Density-Based Counterfactuals for Outliers -- Additional Material

  • 基于局部密度划分空间,实现梯度优化生成反事实
  • 在50个数据集上优于现有方法,逼近度与有效性更优
  • 适合需要解释离群点的工业场景和算法审计

离群点检测旨在识别显著偏离数据主流分布的数据点。解释离群点对理解其成因、验证其重要性及发现潜在偏差或错误至关重要。有效的解释能提供可操作洞察,帮助预防未来类似离群点出现。反事实解释通过识别使特定数据点预测结果改变所需的最小修改,阐明其为何被判定为离群点。尽管价值高,现有反事实解释方法大多忽略离群点检测的独特挑战,且未针对主流的无监督离群点检测算法。局部离群因子(LOF)是最广泛应用的无监督离群点检测方法之一,通过相对局部密度量化离群程度。尽管应用广泛,但其缺乏可解释性。为此,我们提出密度基离群点反事实(DCFO),一种专为LOF设计的新型反事实解释方法。DCFO将数据空间划分为LOF行为平滑的区域,支持高效的梯度优化。在50个OpenML数据集上的大量实验表明,DCFO持续优于基准方法,在生成反事实的接近度和有效性方面表现更优。

原文摘要 · Abstract (English)

Outlier detection identifies data points that significantly deviate from the majority of the data distribution. Explaining outliers is crucial for understanding the underlying factors that contribute to their detection, validating their significance, and identifying potential biases or errors. Effective explanations provide actionable insights, facilitating preventive measures to avoid similar outliers in the future. Counterfactual explanations clarify why specific data points are classified as outliers by identifying minimal changes required to alter their prediction. Although valuable, most existing counterfactual explanation methods overlook the unique challenges posed by outlier detection, and fail to target classical, widely adopted outlier detection algorithms. Local Outlier Factor (LOF) is one the most popular unsupervised outlier detection methods, quantifying outlierness through relative local density. Despite LOF's widespread use across diverse applications, it lacks interpretability. To address this limitation, we introduce Density-based Counterfactuals for Outliers (DCFO), a novel method specifically designed to generate counterfactual explanations for LOF. DCFO partitions the data space into regions where LOF behaves smoothly, enabling efficient gradient-based optimisation. Extensive experimental validation on 50 OpenML datasets demonstrates that DCFO consistently outperforms benchmarked competitors, offering superior proximity and validity of generated counterfactuals.

离群点检测反事实解释可解释性密度方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。