arXiv:2605.29834cs.LG2026-05

用自编码器检测表格数据流中的概念漂移并识别新类别

Open World Autoencoding Drift Detection with Novel Class Recognition in Tabular Non-stationary Data Streams

  • 通过自编码器重构误差检测已知类分布变化
  • 利用代理表示密度估计识别未知新类别样本
  • 支持持续适应动态数据流,适合实时监测场景

数据流处理已成为现代机器学习应用的关键,概念漂移和新类别出现是当前先进识别方法面临的主要挑战。本文提出一种无监督的概念漂移检测方法,基于自编码器的重构误差识别已知类分布的变化,并通过样本代理表示的密度估计实现对新类别样本的识别。采用镜像自编码器结构,使两个任务可独立增量式适应不断变化的数据分布,从而实现对演化概念的持续调整和对未知样本的可靠识别。实验在多种合成表格数据流上进行,同时包含概念漂移与新类别出现。结果表明,该方法在无监督漂移检测和新类别分类方面均达到当前最先进水平。

原文摘要 · Abstract (English)

Data stream processing has become a landmark in modern machine learning applications, with concept drifts and novel class appearances posing the primary challenges faced by sophisticated recognition methods. This work proposes an unsupervised concept drift detection method that identifies shifts in known class distributions based on the reconstruction errors of an autoencoder, while also enabling the recognition of novel class samples through density estimation of a proxy representation of samples. Using mirrored autoencoders allows for independent incremental adaptation to changing problem distributions for the two considered tasks, resulting in continuous adjustment to evolving concepts and reliable recognition of unknown samples. Conducted experiments used a diverse set of synthetic tabular data streams, where both concept drifts and the emergence of novelties were observed. The results show that the proposed approach is competitive with current state-of-the-art unsupervised drift detectors and novelty classifiers.

概念漂移数据流自编码器新类别识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。