用机器学习生成高保真家庭用电合成数据,助力长期电力预测与隐私保护。
Time-series surrogates from energy consumers generated by machine learning approaches for long-term forecasting scenarios
- 融合WGAN、DDPM、HMM和MABF生成个体用电时间序列。
- 合成数据精准复现长期依赖与概率转移特征,提升预测可靠性。
- 兼顾隐私保护,适用于电网规划与状态估计等场景。
电力价值链中的预测研究备受关注,但多数聚焦于发电或用电的短期预测,对个体用户长期用电预测关注不足。本文深入比较多种数据驱动方法,生成适用于长期用电预测的高保真合成时间序列数据。合成数据对电网状态估计与规划至关重要。研究评估了四种前沿但较少使用的技术:混合Wasserstein生成对抗网络(WGAN)、去噪扩散概率模型(DDPM)、隐马尔可夫模型(HMM)和掩码自回归伯恩斯坦多项式归一化流(MABF),分析其在复制个体用电行为的时间动态、长程依赖与概率转移方面的能力。实验基于德国家庭15分钟分辨率的开源数据集,结果揭示各方法优劣,为状态估计等任务提供选型依据。生成框架兼顾数据真实性与隐私保护,防止单个用户被特定识别,所生成的合成用电曲线可直接用于状态估计与消费预测等应用。
原文摘要 · Abstract (English)
Forecasting attracts a lot of research attention in the electricity value chain. However, most studies concentrate on short-term forecasting of generation or consumption with a focus on systems and less on individual consumers. Even more neglected is the topic of long-term forecasting of individual power consumption. Here, we provide an in-depth comparative evaluation of data-driven methods for generating synthetic time series data tailored to energy consumption long-term forecasting. High-fidelity synthetic data is crucial for a wide range of applications, including state estimations in energy systems or power grid planning. In this study, we assess and compare the performance of multiple state-of-the-art but less common techniques: a hybrid Wasserstein Generative Adversarial Network (WGAN), Denoising Diffusion Probabilistic Model (DDPM), Hidden Markov Model (HMM), and Masked Autoregressive Bernstein polynomial normalizing Flows (MABF). We analyze the ability of each method to replicate the temporal dynamics, long-range dependencies, and probabilistic transitions characteristic of individual energy consumption profiles. Our comparative evaluation highlights the strengths and limitations of: WGAN, DDPM, HMM and MABF aiding in selecting the most suitable approach for state estimations and other energy-related tasks. Our generation and analysis framework aims to enhance the accuracy and reliability of synthetic power consumption data while generating data that fulfills criteria like anonymisation - preserving privacy concerns mitigating risks of specific profiling of single customers. This study utilizes an open-source dataset from households in Germany with 15min time resolution. The generated synthetic power profiles can readily be used in applications like state estimations or consumption forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。