跨异构云数据集评测深度异常检测模型表现
Benchmarking Anomaly Detection Across Heterogeneous Cloud Telemetry Datasets
- 统一流程测试4种模型在4类数据上的表现
- 模型性能受特征空间几何与校准稳定性影响显著
- 适合关注云系统异常检测可复现评估的研究者
异常检测对保障云系统可靠性至关重要。尽管深度学习提升了时间序列异常检测能力,但多数模型仅在单一数据集上评估,难以判断其在多类型遥测数据中的泛化能力,尤其在大规模高维环境中。本研究评估了四种深度学习模型(GRU、TCN、Transformer、TSMixer)及经典基线Isolation Forest,在四个不同结构、维度和标注策略的遥测数据集上的表现:Numenta异常基准、微软云监控数据集、Exathlon数据集和IBM控制台数据集。数据集涵盖单变量时间序列、合成多变量负载以及包含超10万特征的真实生产遥测数据。采用统一训练与评估流程,引入类似NAB的指标,支持连续异常区间下的窗口评分,即使标签为点级记录也能实现早期检测行为捕捉。该统一设置使模型行为在相同评分与校准假设下可比。结果表明,云系统异常检测性能不仅取决于模型架构,更关键的是校准稳定性和特征空间几何特性。我们公开预处理流程、基准配置与评估成果,以支持云环境异常检测系统的可复现与部署友好评估。
原文摘要 · Abstract (English)
Anomaly detection is important for keeping cloud systems reliable and stable. Deep learning has improved time-series anomaly detection, but most models are evaluated on one dataset at a time. This raises questions about whether these models can handle different types of telemetry, especially in large-scale and high-dimensional environments. In this study, we evaluate four deep learning models, GRU, TCN, Transformer, and TSMixer. We also include Isolation Forest as a classical baseline. The models are tested across four telemetry datasets: the Numenta Anomaly Benchmark, Microsoft Cloud Monitoring dataset, Exathlon dataset, and IBM Console dataset. These datasets differ in structure, dimensionality, and labelling strategy. They include univariate time series, synthetic multivariate workloads, and real-world production telemetry with over 100,000 features. We use a unified training and evaluation pipeline across all datasets. The evaluation includes NAB-style metrics to capture early detection behaviour for datasets where anomalies persist over contiguous time intervals. This enables window-based scoring in settings where anomalies occur over contiguous time intervals, even when labels are recorded at the point level. The unified setup enables consistent analysis of model behaviour under shared scoring and calibration assumptions. Our results demonstrate that anomaly detection performance in cloud systems is governed not only by model architecture, but critically by calibration stability and feature-space geometry. By releasing our preprocessing pipelines, benchmark configuration, and evaluation artifacts, we aim to support reproducible and deployment-aware evaluation of anomaly detection systems for cloud environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。