平衡异常检测的性能与公平性,提升柴油发电机组监控的可解释性。
Balancing Performance and Fairness in Explainable AI for Anomaly Detection in Distributed Power Plants Monitoring
- 融合集成学习与重采样技术解决极端类别不平衡问题。
- LightGBM达F1-score 0.99,各区域公平性指标DIR≈0.95。
- SHAP揭示油耗和日运行时是关键影响因素,适合运维决策者。
分布式电厂监控中的可靠异常检测对保障运行连续性和降低维护成本至关重要,尤其在依赖柴油发电机的地区。然而,该任务面临极端类别不平衡、可解释性不足及区域间公平性问题。本文提出一种监督学习框架,整合LightGBM、XGBoost等集成模型与支持向量机、K近邻等基线模型,并结合SMOTE与Tomek Links、ENN等先进重采样技术,处理喀麦隆柴油发电机运行数据集中的不平衡问题。通过SHAP实现模型可解释性,使用离散影响比(DIR)量化跨区域公平性,利用最大均值差异(MMD)评估模型泛化能力以捕捉区域间域偏移。实验表明,集成模型持续优于基线,其中LightGBM达到F1-score 0.99,各集群间偏差极小(DIR≈0.95)。SHAP分析显示燃油消耗率和每日运行时长为关键预测因子,为运维人员提供可操作洞察。研究证明可在工业电力管理中同时实现高性能、可解释性与公平性。此外,还讨论了模型在实际中的部署:通过容器化服务实现实时处理、低延迟预测与可解释输出。
原文摘要 · Abstract (English)
Reliable anomaly detection in distributed power plant monitoring systems is essential for ensuring operational continuity and reducing maintenance costs, particularly in regions where telecom operators heavily rely on diesel generators. However, this task is challenged by extreme class imbalance, lack of interpretability, and potential fairness issues across regional clusters. In this work, we propose a supervised ML framework that integrates ensemble methods (LightGBM, XGBoost, Random Forest, CatBoost, GBDT, AdaBoost) and baseline models (Support Vector Machine, K-Nearrest Neighbors, Multilayer Perceptrons, and Logistic Regression) with advanced resampling techniques (SMOTE with Tomek Links and ENN) to address imbalance in a dataset of diesel generator operations in Cameroon. Interpretability is achieved through SHAP (SHapley Additive exPlanations), while fairness is quantified using the Disparate Impact Ratio (DIR) across operational clusters. We further evaluate model generalization using Maximum Mean Discrepancy (MMD) to capture domain shifts between regions. Experimental results show that ensemble models consistently outperform baselines, with LightGBM achieving an F1-score of 0.99 and minimal bias across clusters (DIR $\approx 0.95$). SHAP analysis highlights fuel consumption rate and runtime per day as dominant predictors, providing actionable insights for operators. Our findings demonstrate that it is possible to balance performance, interpretability, and fairness in anomaly detection, paving the way for more equitable and explainable AI systems in industrial power management. {\color{black} Finally, beyond offline evaluation, we also discuss how the trained models can be deployed in practice for real-time monitoring. We show how containerized services can process in real-time, deliver low-latency predictions, and provide interpretable outputs for operators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。