提出标准化框架,让电力保护中的机器学习评估可比可复现
A Standardized Framework for Machine Learning in Power System Protection

- 定义七维评估体系,明确保护目标、观测范围、时间窗口等关键设定
- 在PROTECT-90数据集上,分类F1达0.991,定位误差仅10.20%线路长度
- 强调评估设计需透明,适合研究者和认证机构用于模型可靠性验证
基于机器学习的电力系统保护研究报告的准确率常接近完美,但其真实意义依赖于评估设置。保护任务、物理范围、测量方式、时间窗口、目标定义、预处理及验证协议常混杂变化且描述不全。本文提出一种以标准化为导向的框架,将评估设计视为科学贡献的一部分,定义了七个必需的研究维度:保护目标、物理范围、可观测性、时间与决策窗口、目标与样本有效性、验证协议、评估输出。该框架在公开的PROTECT-90电磁暂态基准数据集上进行有限案例研究,包含9022个模拟事件,来自90 kV双线拓扑结构,用于故障起始条件下的分类与定位。在集中式传感、与仿真元数据对齐的20毫秒窗口、按事件组划分验证条件下,多层感知机(MLP)实现五折平均宏观F1得分为0.991 ± 0.001,定位均方绝对误差为10.20 ± 0.25%线路长度(跨事件组折叠的均值±标准差)。将决策时长扩展至50毫秒后,任务性能不对称性依然存在;可观测性下降使定位误差约翻倍,但对分类影响较小。同步双端传统定位器在更丰富的干净信息下表现优于学习型定位器;测量质量退化表明,干净环境下的预测性能无法决定鲁棒性。该框架将评估假设显式化、可复现,为更可比、可审计的评估及未来面向认证的机器学习保护功能评价奠定基础。
原文摘要 · Abstract (English)
Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends strongly on the evaluation setting. Protection task, physical scope, measurements, timing, targets, preprocessing, and validation often vary jointly and remain incompletely specified. This paper proposes a standardization-oriented framework that treats evaluation design as part of the scientific contribution. It defines seven required study dimensions: protection objective, physical scope, observability, timing and decision windows, targets and sample validity, validation protocol, and evaluation outputs. The framework is instantiated in a bounded case study on the public PROTECT-90 electromagnetic-transient benchmark, comprising 9022 simulated episodes from a 90 kV double-line topology, for onset-conditioned fault classification and localization. Under centralized sensing, simulation-metadata-aligned 20 ms windows, and episode-grouped validation, a multi-layer perceptron (MLP) achieved a five-fold mean macro-averaged F1 score of 0.991 +/- 0.001 for classification and a localization mean absolute error of 10.20 +/- 0.25% of line length (mean +/- std across episode-grouped folds). Extending the decision horizon to 50 ms preserved this task-dependent performance asymmetry, while reduced observability approximately doubled the MLP localization error but had little effect on classification. A synchronized two-ended conventional locator outperformed the learning locators under its richer clean information set, and measurement degradation showed that clean predictive performance did not determine robustness. The framework turns evaluation assumptions into explicit, reproducible evidence and provides a basis for more comparable, auditable evaluation and future certification-oriented assessment of machine-learning protection functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。