arXiv:2608.22968cs.LG2026-08中稿 · poster presentatio…

对比基础模型与轻量级模型在工业监控中的实际效果,发现后者更优。

Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study

论文配图:Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study
图 1 · 摘自论文原文
  • 在三种工业场景中实测对比多种模型性能与成本。
  • 轻量级TCN-AE在异常检测上显著优于基础模型,尤其在鲁棒性上。
  • 尽管基础模型有潜力,但部署成本高且非通用,需按任务选择。

工业监控模型需在满足特定数据、校准和资源约束的前提下,识别操作相关的异常。时间序列基础模型(TSFMs)承诺可复用的表征和零样本预测,但在任务定义多样且轻量级基线表现优异时,其部署价值仍存疑。本研究在三个场景中开展协议感知的实证评估:使用C-MAPSS模拟退化风险,仅用正常数据训练检测异常声音的MIMII数据集,以及带有合成扰动的BDG2预测残差诊断。评估了经典一类分类方法、紧凑神经自编码器、残差预测器、MOMENT-small、Chronos-T5和TimesFM 2.5,指标包括异常排序性能、风险前瞻敏感性、残差预测能力和扰动敏感性,以及本地实现成本。在100个未见的C-MAPSS引擎上,TCN-AE的折合加权AUROC/AUPRC达0.9570/0.8960,远超MOMENT重建的0.7310/0.3080,配对引擎聚类自助法置信区间均排除零。在5个匹配的MIMII泵测试中,OCSVM也优于MOMENT在AUROC和AUPRC上的表现。在固定12米的BDG2面板上,TimesFM 2.5具有最低对齐预测误差和最高的合成AUROC点估计,但合成AUPRC在各模型间相近。同设备测量显示,MOMENT比TCN-AE延迟更高,峰值显存占用更大,状态字典序列化体积也更大。在评估的冻结与零样本设置下,TSFMs是任务依赖的部署选项,而非轻量级模型的默认替代品。

原文摘要 · Abstract (English)

Industrial monitoring models must detect operationally relevant deviations while satisfying target-specific data, calibration, and resource constraints. Time-series foundation models (TSFMs) promise reusable representations and zero-shot forecasts, yet evidence for their deployment value remains mixed when task definitions are heterogeneous and lightweight baselines are competitive. This work presents a protocol-aware empirical assessment across three settings: a C-MAPSS degradation-risk proxy, normal-only training for anomalous-sound detection on MIMII, and BDG2 forecasting-residual diagnostics with synthetic target perturbations. We assess classical one-class methods, compact neural autoencoders, residual forecasters, MOMENT-small, Chronos-T5, and TimesFM 2.5 in terms of anomaly-ranking performance, risk-horizon sensitivity, residual forecasting and perturbation sensitivity, and local implementation cost. Across 100 C-MAPSS engines evaluated out of fold, TCN-AE reaches fold-weighted AUROC/AUPRC 0.9570/0.8960, compared with 0.7310/0.3080 for MOMENT reconstruction; paired engine-cluster bootstrap confidence intervals exclude zero for both differences. Across five matched MIMII pump evaluations, OCSVM also exceeds MOMENT reconstruction in AUROC and AUPRC. On a fixed 12-meter BDG2 panel, TimesFM 2.5 has the lowest aligned forecast error and the highest synthetic AUROC point estimate, although synthetic AUPRC is similar across TSFM and fitted residual models. Same-device measurements show that MOMENT incurs higher latency, peak allocated VRAM, and serialized state-dictionary size than TCN-AE. Under the evaluated frozen and zero-shot settings, TSFMs are task-dependent deployment options rather than default replacements for fitted lightweight models.

工业监控时间序列模型对比成本评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。