arXiv:2512.18977cs.LG2025-12中稿 · Manuscript被引 22

用模糊粗糙集提升异构数据异常检测准确率

Consistency-guided semi-supervised outlier detection in heterogeneous data using fuzzy rough sets

  • 基于标签样本构建模糊相似关系,引导异常检测
  • 结合分类一致性与异常因子,显著降低误报率
  • 适用于含混合类型数据的场景,适合工业异常监控

异常检测旨在识别与多数数据行为不同的样本。半监督方法可利用部分标注信息,降低误报率。然而,现有方法多针对数值型数据,忽视了数据的异质性。本文提出一种基于模糊粗糙集理论的半监督异常检测算法(COD),适用于异构数据。首先,利用少量已知异常样本构建带标签的模糊相似关系;其次,引入模糊决策系统的一致性来评估属性对知识分类的贡献;随后,基于模糊相似类定义异常因子,并融合分类一致性和异常因子进行异常预测。在15个新提出的数据集上进行广泛实验,结果表明COD性能优于或接近当前主流检测器。

原文摘要 · Abstract (English)

Outlier detection aims to find samples that behave differently from the majority of the data. Semi-supervised detection methods can utilize the supervision of partial labels, thus reducing false positive rates. However, most of the current semi-supervised methods focus on numerical data and neglect the heterogeneity of data information. In this paper, we propose a consistency-guided outlier detection algorithm (COD) for heterogeneous data with the fuzzy rough set theory in a semi-supervised manner. First, a few labeled outliers are leveraged to construct label-informed fuzzy similarity relations. Next, the consistency of the fuzzy decision system is introduced to evaluate attributes' contributions to knowledge classification. Subsequently, we define the outlier factor based on the fuzzy similarity class and predict outliers by integrating the classification consistency and the outlier factor. The proposed algorithm is extensively evaluated on 15 freshly proposed datasets. Experimental results demonstrate that COD is better than or comparable with the leading outlier detectors. This manuscript is the accepted author version of a paper published by Elsevier. The final published version is available at https://doi.org/10.1016/j.asoc.2024.112070

异常检测模糊粗糙集半监督异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。