arXiv:2608.23547cs.CRcs.LG2026-08中稿 · and to appear in I…

研究工业控制系统中异常检测模型在训练数据被污染时的鲁棒性,发现模型表现差异大且无法仅凭干净数据性能预测。

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

论文配图:Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination
图 1 · 摘自论文原文
  • 用三种数据污染策略测试11种异构检测器在SWaT数据集上的表现
  • 注入攻击导致局部密度和距离类模型性能下降超50%,特征噪声影响较小
  • 主成分分析、SVM等模型更稳定,适合对数据可信度要求高的工业场景

基于机器学习的异常检测在工业控制系统(ICS)中日益普及,但现有研究多假设训练数据可信。现实中,训练数据可能因日志被篡改、标签错误、历史记录操纵或不安全的重训练过程而被污染。本文在安全水处理(SWaT)基准上评估了离线ICS异常检测流程在训练期数据污染下的鲁棒性。针对11种异构异常检测器,采用三种污染策略:随机注入、相似性目标注入和特征噪声注入。前两种将攻击样本插入正常训练集,第三种在选定正常样本上添加有界高斯噪声。这些污染为基于数据而非梯度的投毒方法。在统一离线协议下,以1%至10%的污染预算,使用干净验证和测试集进行评估。结果表明,鲁棒性显著依赖模型,且无法仅通过干净数据性能预测。基于注入的污染导致最严重退化,尤其影响局部密度和距离类检测器;而特征噪声污染影响较轻。主成分分析(PCA)、支持向量机(SVM)、HBOS和IForest保持相对稳定,而调优后的神经网络检测器表现出中等鲁棒性。总体而言,研究强调了在所评估的数据集、模型与威胁假设下,机器学习驱动的ICS监控必须保障训练数据完整性。

原文摘要 · Abstract (English)

Machine-learning-based anomaly detection is increasingly used in industrial control systems (ICS), yet most studies assume that detector training data is trustworthy. In practice, training data may be corrupted through compromised logs, labeling errors, manipulated historian records, or unsafe retraining processes. This paper evaluates the robustness of offline ICS anomaly-detection pipelines on the Secure Water Treatment (SWaT) benchmark under training-time contamination. We assess 11 heterogeneous anomaly detectors under three contamination strategies: random injection, similarity-targeted injection, and feature-noise injection. The first two insert attack samples into the nominal training pool, while the third adds bounded Gaussian noise to selected normal training samples. These attacks are contamination-based rather than gradient-driven poisoning methods. Contamination budgets from 1% to 10% are evaluated using clean validation and test sets under a unified offline protocol. The results show that robustness is strongly model-dependent and cannot be predicted from clean-data performance alone. Injection-based contamination causes the greatest degradation, particularly for local-density and distance-based detectors, whereas feature-noise contamination has a comparatively limited effect. PCA, SVM, HBOS, and IForest remain relatively stable, while the tuned neural detectors demonstrate intermediate robustness. Overall, the findings highlight the importance of training-data integrity in ML-enabled ICS monitoring, subject to the evaluated dataset, models, and threat assumptions.

异常检测工业控制数据污染鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。