用大模型自动生成并优化电网调度中的安全惩罚函数,减少人工干预。
RL2: Reinforce Large Language Model to Assist Safe Reinforcement Learning for Energy Management of Active Distribution Networks
- 让大模型理解电网安全需求,自动生成惩罚函数。
- 通过多轮对话迭代优化函数,提升强化学习在电网中的安全表现。
- 适合电力系统运维人员和智能调度研究者使用。
随着大规模分布式能源接入主动配电网(ADNs),其能量管理效率较传统配电网显著提升。尽管先进强化学习(RL)方法缓解了复杂建模与优化负担,但安全问题成为实际应用中的关键挑战。由于惩罚函数的设计需依赖大量领域知识,现有方式难以灵活适配不同场景。本文引入大语言模型(LLM)来理解配电网运行安全要求,并自动生成相应惩罚函数。同时提出一种RL2机制,通过多轮对话迭代优化函数的结构与参数,依据下游强化学习代理在训练与测试中的表现动态调整。实验表明,该方法显著降低了对运营人员的依赖,有效提升了安全性与效率。
原文摘要 · Abstract (English)
As large-scale distributed energy resources are integrated into the active distribution networks (ADNs), effective energy management in ADNs becomes increasingly prominent compared to traditional distribution networks. Although advanced reinforcement learning (RL) methods, which alleviate the burden of complicated modelling and optimization, have greatly improved the efficiency of energy management in ADNs, safety becomes a critical concern for RL applications in real-world problems. Since the design and adjustment of penalty functions, which correspond to operational safety constraints, requires extensive domain knowledge in RL and power system operation, the emerging ADN operators call for a more flexible and customized approach to address the penalty functions so that the operational safety and efficiency can be further enhanced. Empowered with strong comprehension, reasoning, and in-context learning capabilities, large language models (LLMs) provide a promising way to assist safe RL for energy management in ADNs. In this paper, we introduce the LLM to comprehend operational safety requirements in ADNs and generate corresponding penalty functions. In addition, we propose an RL2 mechanism to refine the generated functions iteratively and adaptively through multi-round dialogues, in which the LLM agent adjusts the functions' pattern and parameters based on training and test performance of the downstream RL agent. The proposed method significantly reduces the intervention of the ADN operators. Comprehensive test results demonstrate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。