提出新方法解释孤立森林如何判定异常值。
Extending Decision Predicate Graphs for Comprehensive Explanation of Isolation Forest
- 基于决策谓词图构建传播评分,解析异常检测逻辑。
- 可识别关键特征与样本的异常归属依据。
- 适合需透明化异常检测流程的研究者与工程师。
现代机器学习中解释预测模型的需求已得到广泛认可。然而,理解预处理方法同样至关重要。数据变换如何影响模型性能提升及潜在偏差,对构建可靠机器学习流程不可或缺。孤立森林(iForest)是广泛使用的异常检测技术,其效果随树数量增加而提升,但这也使异常值判定与正常值边界解释变得复杂。本文提出一种新型可解释人工智能(XAI)方法,解决全局可解释性问题,旨在为异常检测提供全局解释,克服其黑箱特性。该方法基于决策谓词图(DPG),阐明集成方法的逻辑,并引入内点-外点传播评分(IOP-Score),提供可视化路径与图结构度量,揭示样本被判定为异常的原因。该方法增强iForest可解释性,全面展示决策过程,明确哪些特征参与异常识别及其使用方式。研究推进了异常检测可解释性的前沿水平,提供决策边界洞察与特征整体利用的全景视图,助力实现完全可解释的机器学习流水线。
原文摘要 · Abstract (English)
The need to explain predictive models is well-established in modern machine learning. However, beyond model interpretability, understanding pre-processing methods is equally essential. Understanding how data modifications impact model performance improvements and potential biases and promoting a reliable pipeline is mandatory for developing robust machine learning solutions. Isolation Forest (iForest) is a widely used technique for outlier detection that performs well. Its effectiveness increases with the number of tree-based learners. However, this also complicates the explanation of outlier selection and the decision boundaries for inliers. This research introduces a novel Explainable AI (XAI) method, tackling the problem of global explainability. In detail, it aims to offer a global explanation for outlier detection to address its opaque nature. Our approach is based on the Decision Predicate Graph (DPG), which clarifies the logic of ensemble methods and provides both insights and a graph-based metric to explain how samples are identified as outliers using the proposed Inlier-Outlier Propagation Score (IOP-Score). Our proposal enhances iForest's explainability and provides a comprehensive view of the decision-making process, detailing which features contribute to outlier identification and how the model utilizes them. This method advances the state-of-the-art by providing insights into decision boundaries and a comprehensive view of holistic feature usage in outlier identification. -- thus promoting a fully explainable machine learning pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。