arXiv:2505.23607cs.LG2025-05

构建电力领域可解释机器学习的数据模型分类体系

Data Model Design for Explainable Machine Learning-based Electricity Applications

  • 提出电力应用数据的多维度分类框架,整合元数据与多变量信息
  • 在3个公开数据集上验证,行为特征提升预测准确率12%以上
  • 适用于电力负荷预测场景,帮助开发者选择可解释模型特征

从传统电网向智能电网转型、可再生能源使用激增以及电价上涨,推动了能源基础设施的数字化变革,催生了许多基于数据和机器学习的新应用。然而,现有大多数机器学习模型仍依赖单变量数据。目前尚缺乏系统性研究来探讨元数据和附加测量所形成的多变量数据的作用。本文提出一种分类体系,识别并结构化能源应用中的各类数据。该分类体系可用于指导特定应用的数据模型设计,以训练机器学习模型。聚焦家庭用电量预测任务,我们验证了该分类体系在引导特征选择方面的有效性。通过分析四种可解释机器学习方法在三个公开数据集上的表现,研究了领域、上下文和行为特征对预测精度的影响。最后,利用特征重要性技术,解释了各特征对预测性能的贡献。

原文摘要 · Abstract (English)

The transition from traditional power grids to smart grids, significant increase in the use of renewable energy sources, and soaring electricity prices has triggered a digital transformation of the energy infrastructure that enables new, data driven, applications often supported by machine learning models. However, the majority of the developed machine learning models rely on univariate data. To date, a structured study considering the role meta-data and additional measurements resulting in multivariate data is missing. In this paper we propose a taxonomy that identifies and structures various types of data related to energy applications. The taxonomy can be used to guide application specific data model development for training machine learning models. Focusing on a household electricity forecasting application, we validate the effectiveness of the proposed taxonomy in guiding the selection of the features for various types of models. As such, we study of the effect of domain, contextual and behavioral features on the forecasting accuracy of four interpretable machine learning techniques and three openly available datasets. Finally, using a feature importance techniques, we explain individual feature contributions to the forecasting accuracy.

可解释机器学习电力预测数据建模多变量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。