用监督强化学习提升分布式能源协调效率,降低训练成本。
Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources

- 先用示范数据监督预训练,再用强化学习微调策略
- 两阶段微调使策略在低质量数据下仍具高成本效益
- 适合电力系统优化与智能调度方向的研究者
分布式能源资源(DERs)的规模化接入对电力系统脱碳至关重要,但其固有的不确定性与建模复杂性制约了灵活性释放。传统优化方法难以应对此类不确定性与复杂性,强化学习(RL)成为一种有前景的替代方案。然而,标准RL方法在从零训练时存在样本效率低、性能不足的问题。受大语言模型训练范式的启发,本文提出一种监督强化学习(SRL)框架,用于学习DER协调策略:首先以监督学习方式在示范数据上预训练策略,随后通过强化学习进行微调。进一步设计两步微调机制——离线微调以提升策略性能,在线微调以适应真实动态环境。实验表明,基于该框架的RL实现显著优于所有基准,在低质量示范数据下仍能保持高成本效益。
原文摘要 · Abstract (English)
The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by their inherent uncertainties and modelling complexity. As traditional optimization methods struggle with such uncertainty and complexity of DERs, reinforcement learning (RL) has emerged as a promising alternative for DER management. However, standard RL methods suffer from sample inefficiency and sub-optimality when trained from scratch. Inspired by the training paradigms in large language models, this paper proposes a Supervised Reinforcement Learning (SRL) framework for learning DER coordination policies. This framework first pre-trains a policy on demonstration data in a supervised-learning fashion, which is then further fine-tuned using RL. Furthermore, we propose a two-step fine-tuning process: offline fine-tuning for enhancing policy performance and online fine-tuning for adapting it to the real-world dynamics. Experiments demonstrate that RL implementations based on the proposed framework significantly outperform all benchmarks, achieving high cost efficiency even under low-quality demonstration data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。