用深度学习预测电价,让电池智能买卖电赚更多钱。
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
- 结合深度强化学习与多时序预测,让电池提前规划充放电。
- 在加拿大阿尔伯塔省数据上,收益比无预测提升60%。
- 多个不完美预测组合能提升决策效果,适合电力交易研究者。
储能能量套利是电池运营商最盈利的收入来源之一,通过在不同电价时买入和卖出电能获利。由于电价固有的不确定性,预测收益极具挑战性。近年来,深度强化学习(DRL)因其可基于大量历史数据处理不确定性而成为有力工具。然而,因无法获取未来电价,传统DRL代理只能响应当前价格,难以学习前瞻性的电池调度策略。为此,本研究将DRL与深度学习中的时间序列预测方法相结合,以提升能量套利性能。我们在加拿大阿尔伯塔省的电价数据上开展案例研究,该数据以价格剧烈波动和高度非平稳性为特征,即便采用包含卷积层、循环层和注意力模块的先进深度学习模型也难以准确预测。结果表明,即使预测不完美,基于DRL的电池控制仍显著受益于这些预测,前提是整合多个预测窗口。将未来24小时内的多个预测结果进行分组后,深度Q网络(DQN)的累积奖励相比无预测情形提升了60%。我们推测,尽管各预测存在误差,但通过‘多数表决’机制,它们共同传递了未来电价走势的有用信息,从而帮助DRL代理学习更优的控制策略。
原文摘要 · Abstract (English)
Energy arbitrage is one of the most profitable sources of income for battery operators, generating revenues by buying and selling electricity at different prices. Forecasting these revenues is challenging due to the inherent uncertainty of electricity prices. Deep reinforcement learning (DRL) emerged in recent years as a promising tool, able to cope with uncertainty by training on large quantities of historical data. However, without access to future electricity prices, DRL agents can only react to the currently observed price and not learn to plan battery dispatch. Therefore, in this study, we combine DRL with time-series forecasting methods from deep learning to enhance the performance on energy arbitrage. We conduct a case study using price data from Alberta, Canada that is characterized by irregular price spikes and highly non-stationary. This data is challenging to forecast even when state-of-the-art deep learning models consisting of convolutional layers, recurrent layers, and attention modules are deployed. Our results show that energy arbitrage with DRL-enabled battery control still significantly benefits from these imperfect predictions, but only if predictors for several horizons are combined. Grouping multiple predictions for the next 24-hour window, accumulated rewards increased by 60% for deep Q-networks (DQN) compared to the experiments without forecasts. We hypothesize that multiple predictors, despite their imperfections, convey useful information regarding the future development of electricity prices through a "majority vote" principle, enabling the DRL agent to learn more profitable control policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。