arXiv:2510.18998cs.LGcs.DB2025-10中稿 · ICDE 2026被引 4

用编码后分解方法提升异常检测在污染数据下的鲁棒性

An Encode-then-Decompose Approach to Unsupervised Time Series Anomaly Detection on Contaminated Training Data--Extended Version

  • 先编码再分解表示,分离出稳定与辅助成分
  • 在8个基准上表现达顶尖水平,污染率变化下仍稳定
  • 用互信息替代重构误差,更适配含噪训练数据

时间序列异常检测在现代大规模系统中至关重要,广泛应用于各类系统监控。无监督方法因无需训练时的异常标签而受到广泛关注,降低了标注成本并拓展了应用范围。其中,自编码器因其利用压缩表示的重构误差定义异常分数而备受关注。然而,自编码器学习的表示对训练数据中的异常敏感,导致性能下降。本文提出一种新的编码-分解范式,将编码后的表示分解为稳定和辅助两部分,从而增强在含污染时间序列训练下的鲁棒性。此外,我们引入一种基于互信息的新度量指标,取代传统的重构误差来识别异常。所提方法在八个常用多变量与单变量时间序列基准上表现出竞争性或领先性能,并对不同污染比例的时间序列具有强鲁棒性。

原文摘要 · Abstract (English)

Time series anomaly detection is important in modern large-scale systems and is applied in a variety of domains to analyze and monitor the operation of diverse systems. Unsupervised approaches have received widespread interest, as they do not require anomaly labels during training, thus avoiding potentially high costs and having wider applications. Among these, autoencoders have received extensive attention. They use reconstruction errors from compressed representations to define anomaly scores. However, representations learned by autoencoders are sensitive to anomalies in training time series, causing reduced accuracy. We propose a novel encode-then-decompose paradigm, where we decompose the encoded representation into stable and auxiliary representations, thereby enhancing the robustness when training with contaminated time series. In addition, we propose a novel mutual information based metric to replace the reconstruction errors for identifying anomalies. Our proposal demonstrates competitive or state-of-the-art performance on eight commonly used multi- and univariate time series benchmarks and exhibits robustness to time series with different contamination ratios.

异常检测自编码器时间序列鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。