arXiv:2504.04326eess.SYcs.LG2025-04被引 4

用简单规则生成示范数据,提升电池储能长期经济调度的强化学习效果

Economic Battery Storage Dispatch with Deep Reinforcement Learning from Rule-Based Demonstrations

  • 基于电价规则生成示范数据,辅助深度强化学习训练
  • 样本效率和最终收益显著提升,年周期调度表现更优
  • 对规则选择不敏感,适合电力系统优化初学者使用

深度强化学习在经济型电池调度中的应用近年显著增长,但长周期调度因奖励延迟导致性能不佳。实验发现,采用小时级分辨率的年度任务中,主流演员-评论家算法表现较差。为此,本文提出一种扩展软演员-评论家(SAC)的方法,通过规则基础策略生成示范数据,替代专家示范。我们以电网连接微电网为案例,利用基于现货电价的if-then-else规则生成示范数据,并存入独立回放缓冲区,按线性衰减概率与智能体自身经验混合采样。尽管示范数据粗糙,改进仍显著提升样本效率与最终回报。进一步表明,该方法稳定优于示范规则,且对规则选择具有鲁棒性,只要规则能引导初期训练方向即可。

原文摘要 · Abstract (English)

The application of deep reinforcement learning algorithms to economic battery dispatch problems has significantly increased recently. However, optimizing battery dispatch over long horizons can be challenging due to delayed rewards. In our experiments we observe poor performance of popular actor-critic algorithms when trained on yearly episodes with hourly resolution. To address this, we propose an approach extending soft actor-critic (SAC) with learning from demonstrations. The special feature of our approach is that, due to the absence of expert demonstrations, the demonstration data is generated through simple, rule-based policies. We conduct a case study on a grid-connected microgrid and use if-then-else statements based on the wholesale price of electricity to collect demonstrations. These are stored in a separate replay buffer and sampled with linearly decaying probability along with the agent's own experiences. Despite these minimal modifications and the imperfections in the demonstration data, the results show a drastic performance improvement regarding both sample efficiency and final rewards. We further show that the proposed method reliably outperforms the demonstrator and is robust to the choice of rule, as long as the rule is sufficient to guide early training into the right direction.

强化学习电池调度能源管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。