构建最大多变量时间序列异常检测基准,揭示模型选择仍需突破。
mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
- 构建344个带标签数据集的跨领域基准,覆盖19个真实场景。
- 24种检测器表现各异,无单一方法在所有数据上领先。
- 现有模型选择方法仍不理想,亟需更鲁棒的通用策略。
多变量时间序列异常检测在医疗、网络安全和工业监控等领域至关重要,但因高维依赖、变量间时序相关性及标注异常稀缺而极具挑战。我们提出mTSBench,目前规模最大的多变量时间序列异常检测与模型选择基准,包含来自19个应用领域的344个带标签时间序列。我们全面评估了24种异常检测器,包括仅有的两种公开可用的基于大语言模型的方法。结果表明,无单一检测器在所有数据集上占优,凸显有效模型选择的必要性。我们还测试了三种近期模型选择方法,发现即使最强者也远未达到最优。研究强调了对鲁棒、可泛化选择策略的迫切需求。基准已开源,地址为https://plan-lab.github.io/mtsbench,以推动未来研究。
原文摘要 · Abstract (English)
Anomaly detection in multivariate time series is essential across domains such as healthcare, cybersecurity, and industrial monitoring, yet remains fundamentally challenging due to high-dimensional dependencies, the presence of cross-correlations between time-dependent variables, and the scarcity of labeled anomalies. We introduce mTSBench, the largest benchmark to date for multivariate time series anomaly detection and model selection, consisting of 344 labeled time series across 19 datasets from a wide range of application domains. We comprehensively evaluate 24 anomaly detectors, including the only two publicly available large language model-based methods for multivariate time series. Consistent with prior findings, we observe that no single detector dominates across datasets, motivating the need for effective model selection. We benchmark three recent model selection methods and find that even the strongest of them remain far from optimal. Our results highlight the outstanding need for robust, generalizable selection strategies. We open-source the benchmark at https://plan-lab.github.io/mtsbench to encourage future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。