arXiv:2506.06484eess.SYcs.AI2025-06中稿 · publication at the…被引 1

用深度强化学习优化储能型电转气系统,解决长期收益延迟难题

The Economic Dispatch of Power-to-Gas Systems with Deep Reinforcement Learning:Tackling the Challenge of Delayed Rewards with Long-Term Energy Storage

  • 结合预测与奖励惩罚机制改进DRL算法,应对能源转换延迟问题
  • 在多场景测试中显著提升系统运行成本控制能力,降低3.2%~8.7%综合支出
  • 适合研究能源系统调度、智能电网优化的工程师与研究人员参考

电转气(Power-to-Gas, P2G)技术因其可促进风能、太阳能等间歇性可再生能源并网而受到关注。然而,受可再生能源波动性、电价与负荷变化影响,其经济运行优化极为复杂。相比电池储能系统(BES),P2G能量转化与存储效率较低,且效益显现周期长。深度强化学习(DRL)虽在处理不确定性方面有潜力,但面临决策延迟回报的挑战。以往研究多聚焦短期转换过程,忽视了P2G的长期储能能力。本研究系统评估DRL算法(如Deep Q-Network和Proximal Policy Optimization)在含BES与燃气轮机的长期运行场景中的表现,并提出三项改进:引入预测信息、在奖励函数中加入惩罚项、采用策略性成本计算。通过三个逐步复杂的案例研究发现,尽管初始阶段DRL难以应对复杂决策,但经上述调整后,其制定低成本运行策略的能力显著提升,有效释放了P2G长期储能的潜力。

原文摘要 · Abstract (English)

Power-to-Gas (P2G) technologies gain recognition for enabling the integration of intermittent renewables, such as wind and solar, into electricity grids. However, determining the most cost-effective operation of these systems is complex due to the volatile nature of renewable energy, electricity prices, and loads. Additionally, P2G systems are less efficient in converting and storing energy compared to battery energy storage systems (BESs), and the benefits of converting electricity into gas are not immediately apparent. Deep Reinforcement Learning (DRL) has shown promise in managing the operation of energy systems amidst these uncertainties. Yet, DRL techniques face difficulties with the delayed reward characteristic of P2G system operation. Previous research has mostly focused on short-term studies that look at the energy conversion process, neglecting the long-term storage capabilities of P2G. This study presents a new method by thoroughly examining how DRL can be applied to the economic operation of P2G systems, in combination with BESs and gas turbines, over extended periods. Through three progressively more complex case studies, we assess the performance of DRL algorithms, specifically Deep Q-Networks and Proximal Policy Optimization, and introduce modifications to enhance their effectiveness. These modifications include integrating forecasts, implementing penalties on the reward function, and applying strategic cost calculations, all aimed at addressing the issue of delayed rewards. Our findings indicate that while DRL initially struggles with the complex decision-making required for P2G system operation, the adjustments we propose significantly improve its capability to devise cost-effective operation strategies, thereby unlocking the potential for long-term energy storage in P2G technologies.

电转气强化学习能源调度长期储能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。