arXiv:2602.15457cs.LG2026-02

提出事件级评估框架,更真实地测试物联网时序异常检测模型性能。

Benchmarking IoT Time-Series AD with Event-Level Augmentations

  • 统一引入事件级扰动模拟真实场景问题
  • 不同模型在噪声和漂移下表现差异显著,无通用最优方案
  • 适合工业界选型参考与模型可解释性分析

针对安全关键的物联网时序异常检测,应以事件级别评估可靠性与早期发现能力,而非仅依赖精心筛选数据集上的点级结果。本文提出统一的事件级增强评估协议,模拟校准传感器失效、线性与对数漂移、加性噪声及窗口偏移等现实问题。通过通道级掩码缺失与影响估计实现传感器层面探查,支持根因分析。在五个公开数据集(SWaT、WADI、SMD、SKAB、TEP)和两个工业数据集(蒸汽轮机、核涡轮发电机)上,使用统一划分与事件聚合评估14个代表性模型。结果显示:图结构模型在丢包和长事件中迁移性最好(如在SWaT上加噪时图自编码器F1从0.804降至0.677);密度/流模型在平稳工况下表现优但对单调漂移敏感;谱卷积网络在强周期性场景领先;重构自编码器经基础传感器筛选后竞争力提升;预测/混合动态模型在破坏时序依赖故障时有效但对窗口敏感。评估还揭示设计选择影响:在SWaT对数漂移下,用高斯密度替代归一化流使高压应力下F1从约0.75降至约0.57;固定学习的有向无环图带来约0.5–1.0分清洁集增益,但漂移敏感度增加约8倍。

原文摘要 · Abstract (English)

Anomaly detection (AD) for safety-critical IoT time series should be judged at the event level: reliability and earliness under realistic perturbations. Yet many studies still emphasize point-level results on curated base datasets, limiting value for model selection in practice. We introduce an evaluation protocol with unified event-level augmentations that simulate real-world issues: calibrated sensor dropout, linear and log drift, additive noise, and window shifts. We also perform sensor-level probing via mask-as-missing zeroing with per-channel influence estimation to support root-cause analysis. We evaluate 14 representative models on five public anomaly datasets (SWaT, WADI, SMD, SKAB, TEP) and two industrial datasets (steam turbine, nuclear turbogenerator) using unified splits and event aggregation. There is no universal winner: graph-structured models transfer best under dropout and long events (e.g., on SWaT under additive noise F1 drops 0.804->0.677 for a graph autoencoder, 0.759->0.680 for a graph-attention variant, and 0.762->0.756 for a hybrid graph attention model); density/flow models work well on clean stationary plants but can be fragile to monotone drift; spectral CNNs lead when periodicity is strong; reconstruction autoencoders become competitive after basic sensor vetting; predictive/hybrid dynamics help when faults break temporal dependencies but remain window-sensitive. The protocol also informs design choices: on SWaT under log drift, replacing normalizing flows with Gaussian density reduces high-stress F1 from ~0.75 to ~0.57, and fixing a learned DAG gives a small clean-set gain (~0.5-1.0 points) but increases drift sensitivity by ~8x.

异常检测物联网事件级评估工业时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。