用聚类算法自动筛选风场SCADA数据中的异常,提升效率与准确性。
Clustering algorithms for multivariate wind farm SCADA data filtering

- 基于多变量10分钟统计特征,用聚类方法识别正常运行数据。
- 聚类方法检测出明显和隐蔽异常,准确率高于人工视觉筛选。
- 适合风电运维人员、数据工程师用于自动化数据清洗。
风场运行中,监控与数据采集(SCADA)系统记录大量异常、瞬态及特定工况数据,形成海量数据集。但多数应用仅需正常运行数据,因此必须对数据进行过滤。为此,本文比较多种聚类算法在自动过滤方面的准确性,并提出适用于无标签数据且鲁棒性强的评估指标。研究基于某海上风电场3台风机的SCADA数据,采用多通道10分钟统计特征进行分析。除常规异常与运行模式外,数据还包含大量因现场测试产生的非明显离群值。结果表明,扩展分析维度(超越功率曲线)在特征选择与评估指标设计中至关重要。多数聚类方法能有效识别明显与细微异常,准确率优于人工筛选;但不同模型的准确率与保留数据量差异显著,仍需专家适度介入,相比传统人工方式大幅减少工作量。
原文摘要 · Abstract (English)
During wind farm operation, Supervisory Control and Data Acquisition (SCADA) systems record numerous anomalies, transients, and specific operational modes, leading to large datasets. However, for a wide range of applications, only measurements corresponding to normal operation are required and, therefore, the SCADA data must be filtered. For this purpose, several methods have been proposed to automate and replace manual filtering conducted by experts via visual inspection of the data. In this paper, we compare the filtering accuracy of multiple clustering algorithms against manual filtering, introducing evaluation metrics that are suitable for unlabeled data and robust across potential applications. Based on the results, we provide recommendations for generalizing model calibration to different datasets and discuss potential use cases for each model. The models are applied to the SCADA data of three turbines of an existing offshore wind farm, using 10-minute statistics across multiple data channels. In addition to the anomalies and operational modes typically recorded, the dataset presents a large number of non-evident outliers due to several field tests. Overall, the results highlight the importance of extending the analysis beyond the power curve, both in feature selection and in the design of evaluation metrics. In most cases, cluster-based methods are able to detect both evident and subtle outliers, achieving higher accuracy than manual filtering. However, the accuracy and the amount of data retained vary considerably depending on the model, and expert involvement remains necessary, though to a reduced extent compared to manual filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。