用解释方法分析异常检测模型差异,选互补的模型组成更强集成
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
- 用SHAP分析模型对输入特征的重要性分配,量化决策机制
- 解释差异大的模型产生不重叠的异常结果,更互补
- 强调模型性能与解释多样性并重,提升集成效果
无监督异常检测因数据分布多样且缺乏标签而具有挑战性。集成方法通过组合多个检测器来缓解这些问题,但许多检测器依赖相似决策线索,导致异常评分冗余。为解决此问题,我们提出一种基于决策机制表征检测器的方法。利用SHapley Additive exPlanations(SHAP)量化各模型对输入特征的重要性分配,并通过这些归因谱图衡量检测器间的相似性。结果显示,具有相似解释的检测器其异常评分高度相关,且识别出的异常大量重叠;而解释差异大则可靠指示互补检测行为。实验表明,基于解释的度量可提供不同于原始输出的新筛选标准。然而我们也发现,仅追求多样性不足;高质量单个模型仍是有效集成的前提。通过同时优化解释多样性与模型质量,我们构建的集成在多样性、互补性及整体性能上均更优。
原文摘要 · Abstract (English)
Unsupervised anomaly detection is a challenging problem due to the diversity of data distributions and the lack of labels. Ensemble methods are often adopted to mitigate these challenges by combining multiple detectors, which can reduce individual biases and increase robustness. Yet building an ensemble that is genuinely complementary remains challenging, since many detectors rely on similar decision cues and end up producing redundant anomaly scores. As a result, the potential of ensemble learning is often limited by the difficulty of identifying models that truly capture different types of irregularities. To address this, we propose a methodology for characterizing anomaly detectors through their decision mechanisms. Using SHapley Additive exPlanations, we quantify how each model attributes importance to input features, and we use these attribution profiles to measure similarity between detectors. We show that detectors with similar explanations tend to produce correlated anomaly scores and identify largely overlapping anomalies. Conversely, explanation divergence reliably indicates complementary detection behavior. Our results demonstrate that explanation-driven metrics offer a different criterion than raw outputs for selecting models in an ensemble. However, we also demonstrate that diversity alone is insufficient; high individual model performance remains a prerequisite for effective ensembles. By explicitly targeting explanation diversity while maintaining model quality, we are able to construct ensembles that are more diverse, more complementary, and ultimately more effective for unsupervised anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。