用模糊近似与相对熵结合,提升混合属性数据的异常检测效果
Outlier detection in mixed-attribute data: a semi-supervised approach with fuzzy approximations and relative entropy
- 基于模糊近似评估属性集对异常检测的贡献
- 通过模糊相对熵量化不确定性,识别偏离正常模式的数据
- 在16个公开数据集上优于或媲美主流算法,适合处理复杂真实数据
异常检测是数据挖掘中的关键任务,旨在识别显著偏离正常模式的对象。半监督方法通过利用部分标记数据提升检测性能,但通常忽视真实混合属性数据中的不确定性和异质性。本文提出一种基于模糊粗糙集的异常检测方法(FROD),首先使用少量标记数据构建模糊决策系统,通过模糊近似计算属性分类准确率,评估属性集对异常检测的贡献;随后利用未标记数据计算模糊相对熵,从不确定性角度刻画异常;最后结合两者设计检测算法。在16个公开数据集上的实验表明,FROD性能可与或优于主流算法。所有数据集与源码可在https://github.com/ChenBaiyang/FROD获取。
原文摘要 · Abstract (English)
Outlier detection is a critical task in data mining, aimed at identifying objects that significantly deviate from the norm. Semi-supervised methods improve detection performance by leveraging partially labeled data but typically overlook the uncertainty and heterogeneity of real-world mixed-attribute data. This paper introduces a semi-supervised outlier detection method, namely fuzzy rough sets-based outlier detection (FROD), to effectively handle these challenges. Specifically, we first utilize a small subset of labeled data to construct fuzzy decision systems, through which we introduce the attribute classification accuracy based on fuzzy approximations to evaluate the contribution of attribute sets in outlier detection. Unlabeled data is then used to compute fuzzy relative entropy, which provides a characterization of outliers from the perspective of uncertainty. Finally, we develop the detection algorithm by combining attribute classification accuracy with fuzzy relative entropy. Experimental results on 16 public datasets show that FROD is comparable with or better than leading detection algorithms. All datasets and source codes are accessible at https://github.com/ChenBaiyang/FROD. This manuscript is the accepted author version of a paper published by Elsevier. The final published version is available at https://doi.org/10.1016/j.ijar.2025.109373
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。