简单PCA模型在特定条件下可媲美复杂深度学习模型
Revisiting OmniAnomaly for Anomaly Detection: performance metrics and comparison with PCA-based models
- 用相同评估流程对比OmniAnomaly与PCA模型
- PCA在无点调整时表现优于甚至等于OmniAnomaly
- 强调评估方法比模型复杂度更影响结果可信度
深度学习模型已成为多变量时间序列异常检测(MTSAD)的主流方法,常宣称显著优于传统统计方法。但这些性能提升多在异构阈值策略和评估协议下进行,导致公平比较困难。本文重新审视广泛使用的随机循环模型OmniAnomaly,将其与基于主成分分析(PCA)的简单线性基线在服务器机器数据集(SMD)上进行系统对比。两个方法均采用相同的阈值设定与评估流程,每个28台机器重复100次实验。在点级别使用精确率、召回率和F1分数评估,考虑点调整与否及跨机器与实验的聚合策略,并报告对应标准差。结果表明各机器间性能差异显著,且在未应用点调整时,PCA性能可媲美甚至超越OmniAnomaly。该发现质疑了当前基准测试实践中复杂架构的实际价值,凸显评估方法在MTSAD研究中的关键作用。
原文摘要 · Abstract (English)
Deep learning models have become the dominant approach for multivariate time series anomaly detection (MTSAD), often reporting substantial performance improvements over classical statistical methods. However, these gains are frequently evaluated under heterogeneous thresholding strategies and evaluation protocols, making fair comparisons difficult. This work revisits OmniAnomaly, a widely used stochastic recurrent model for MTSAD, and systematically compares it with a simple linear baseline based on Principal Component Analysis (PCA) on the Server Machine Dataset (SMD). Both methods are evaluated under identical thresholding and evaluation procedures, with experiments repeated across 100 runs for each of the 28 machines in the dataset. Performance is evaluated using Precision, Recall and F1-score at point-level, with and without point-adjustment, and under different aggregation strategies across machines and runs, with the corresponding standard deviations also reported. The results show large variability across machines and show that PCA can achieve performance comparable to OmniAnomaly, and even outperform it when point-adjustment is not applied. These findings question the added value of more complex architectures under current benchmarking practices and highlight the critical role of evaluation methodology in MTSAD research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。