arXiv:2605.10822cs.LGeess.SP2026-05

评测预测模型在传感器故障下的鲁棒性,发现干净数据表现好的模型可能在故障下严重失效。

Benchmarking Sensor-Fault Robustness in Forecasting

论文配图:Benchmarking Sensor-Fault Robustness in Forecasting
图 1 · 摘自论文原文
  • 构建了基于真实工业数据的传感器故障测试基准,统一评估各类模型鲁棒性。
  • 实测显示,部分模型在故障场景下性能下降超50%,清洁数据排名与故障排名不一致。
  • 适合关注工业预测系统可靠性的研究者和工程师使用。

网络物理系统(CPS)的预测模型依赖于含噪声、偏差、缺失或时间错位的传感器数据,但常规评估常仅以无故障误差为标准,未考察模型在实际故障下的表现。本文提出SensorFault-Bench——一个面向真实工业数据的标准化传感器故障压力测试协议,用于评估预测架构与鲁棒性改进方法,并建立可操作的分类体系。在四个真实数据集上,通过八个评分场景,依据统一的严重度模型,报告最差场景下的退化程度、清洁均方误差(MSE)以及故障时段的均方误差,分离出相对鲁棒性与绝对误差。采用独立故障迁移划分,使显式训练方法在相邻故障类型上训练,评估则使用独立基准场景。实验表明,以清洁数据误差优胜的模型在故障下可能急剧退化,且清洁误差排名与最差场景故障时误差排名不一致。被评估的零样本基础模型Chronos-2在两个单目标数据集上,其清洁MSE甚至不如简单后向值预测器,在ETTh1和Traffic数据集上的最差场景退化幅度最大。对于鲁棒性改进方法,投影梯度下降对抗训练与随机训练在值类故障主导场景中有效降低退化,而故障增强策略在可用性故障主导场景中更优。该基准提供开源代码、文档化数据访问及复现扩展指南,支持新数据集、架构和方法在相同传感器故障鲁棒性协议下进行评估。

原文摘要 · Abstract (English)

Cyber-physical system (CPS) forecasting models depend on sensor streams with noisy, biased, missing, or temporally misaligned readings, yet standard forecasting evaluation often selects models by nominal error without showing whether they remain robust under such faults. We introduce SensorFault-Bench, a shared CPS-grounded sensor-fault stress-test protocol for evaluating forecasting architectures and robustness-improvement methods, and an operational taxonomy organizing the method comparison. Across four real-world datasets and eight scored scenarios governed by a standardized severity model, it reports worst-scenario degradation, clean mean squared error (MSE), and worst-scenario fault-time MSE, separating relative robustness from absolute error. A disjoint fault-transfer split lets explicit fault-training methods train on adjacent fault families while evaluation uses separate benchmark scenarios. Empirically, forecasting architectures favored by clean MSE can degrade sharply under faults, and clean-MSE rankings can disagree with worst-scenario fault-time error rankings. Chronos-2, the evaluated zero-shot foundation-model representative, matches or trails the last-value naive forecaster in clean MSE on the two single-target datasets and has the largest worst-scenario degradation on ETTh1 and Traffic, where all channels are forecast targets. For the evaluated robustness-improvement method set, paired deltas show selective degradation reductions: projected gradient descent adversarial training and randomized training lead where value faults dominate observed degradation, while fault augmentation leads where availability faults dominate. SensorFault-Bench provides open-source code, documented data access, and reproduction and extension guides, so new datasets, architectures, and robustness-improvement methods can be evaluated under the same CPS sensor-fault robustness protocol.

预测模型传感器故障鲁棒性评测工业系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。