工业时序数据复杂度高,自编码器比传统方法更有效检测异常。
Unsupervised Anomaly Detection in Process-Complex Industrial Time Series: A Real-World Case Study

- 用自编码器建模工业时序的多尺度非周期动态
- 时序卷积自编码器在真实数据上表现最稳定
- 适合工业界部署的异常检测研究者参考
来自真实生产环境的工业时序数据远比常用基准数据集复杂,主要源于异构、多阶段运行流程。因此,在简化条件下验证的方法常无法推广到工业场景。本文基于一套从全运行状态工业设备采集的独特数据集,系统评估不同模型对过程引发的显著变异性建模能力。从经典的孤立森林基线出发,扩展至多种自编码器架构。实验表明,孤立森林难以捕捉数据中的非周期、多尺度动态,而自编码器整体表现更优;其中,时序卷积自编码器性能最稳健,递归与变分变体则需更精细调参。
原文摘要 · Abstract (English)
Industrial time-series data from real production environments exhibits substantially higher complexity than commonly used benchmark datasets, primarily due to heterogeneous, multi-stage operational processes. As a result, anomaly detection methods validated under simplified conditions often fail to generalize to industrial settings. This work presents an empirical study on a unique dataset collected from fully operational industrial machinery, explicitly capturing pronounced process-induced variability. We evaluate which model classes are capable of capturing this complexity, starting with a classical Isolation Forest baseline and extending to multiple autoencoder architectures. Experimental results show that Isolation Forest is insufficient for modeling the non-periodic, multi-scale dynamics present in the data, whereas autoencoders consistently perform better. Among them, temporal convolutional autoencoders achieve the most robust performance, while recurrent and variational variants require more careful tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。