arXiv:2411.13951cs.LGcs.AI2024-11被引 2

构建首个用于多变量时间序列在线无监督异常检测的离散序列数据集。

PATH: A Discrete-sequence Dataset for Evaluating Online Unsupervised Anomaly Detection Approaches for Multivariate Time Series

  • 基于先进仿真工具生成真实汽车动力系统数据,涵盖多变量动态特性。
  • 提供含污染与纯净版本的训练测试集,支持无监督与半监督设置。
  • 揭示阈值选择对检测性能影响大,推动无需标注数据的自适应方法研究。

多变量时间序列的异常检测基准评估面临高质量数据集匮乏的挑战。现有公开数据集规模小、多样性差且异常过于简单,阻碍了该领域的可衡量进展。本文提出一种通过前沿仿真工具生成的多样化、大规模、非平凡数据集,模拟真实汽车动力系统的多变量、动态及变状态行为。此外,该数据集为离散序列问题,此前文献未涉及。为适配无监督与半监督异常检测、时间序列生成与预测任务,我们提供多种版本数据集,其中训练与测试子集分别以污染和纯净形式提供。我们还提供了基于确定性与变分自编码器及非参数方法的基线结果。实验表明,基于半监督版本训练的方法优于无监督方法,凸显对训练数据污染更具鲁棒性的算法需求。同时,阈值选择显著影响检测性能,提示需更多研究探索无需标签数据的自适应阈值方法。

原文摘要 · Abstract (English)

Benchmarking anomaly detection approaches for multivariate time series is a challenging task due to a lack of high-quality datasets. Current publicly available datasets are too small, not diverse and feature trivial anomalies, which hinders measurable progress in this research area. We propose a solution: a diverse, extensive, and non-trivial dataset generated via state-of-the-art simulation tools that reflects realistic behaviour of an automotive powertrain, including its multivariate, dynamic and variable-state properties. Additionally, our dataset represents a discrete-sequence problem, which remains unaddressed by previously-proposed solutions in literature. To cater for both unsupervised and semi-supervised anomaly detection settings, as well as time series generation and forecasting, we make different versions of the dataset available, where training and test subsets are offered in contaminated and clean versions, depending on the task. We also provide baseline results from a selection of approaches based on deterministic and variational autoencoders, as well as a non-parametric approach. As expected, the baseline experimentation shows that the approaches trained on the semi-supervised version of the dataset outperform their unsupervised counterparts, highlighting a need for approaches more robust to contaminated training data. Furthermore, results show that the threshold used can have a large influence on detection performance, hence more work needs to be invested in methods to find a suitable threshold without the need for labelled data.

异常检测时间序列数据集汽车系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。