用多目标优化自动设计强化学习环境,提升电力系统最优潮流求解效果
A General Approach of Automated Environment Design for Learning the Optimal Power Flow
- 基于超参数优化框架,自动配置强化学习环境参数
- 在5个基准问题上均超越人工设计环境,训练性能更优
- 揭示关键设计因素,避免算法过拟合,适合电网优化研究者
强化学习(RL)算法被越来越多地用于求解最优电力潮流(OPF)问题。然而,如何设计强化学习环境以最大化训练性能,无论是针对OPF还是通用场景,仍无定论。本文提出一种通用的自动化强化学习环境设计方法,利用多目标优化实现。该方法基于超参数优化(HPO)框架,可复用现有HPO算法与技术。在五个OPF基准问题上,我们证明所提出的自动化设计方法始终优于人工构建的基线环境。此外,通过统计分析确定了影响性能的关键环境设计决策,获得多项关于强化学习-最优潮流环境设计的新见解。最后,讨论了环境对所用强化学习算法过拟合的风险。据我们所知,这是首个通用的自动化强化学习环境设计方法。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms are increasingly used to solve the optimal power flow (OPF) problem. Yet, the question of how to design RL environments to maximize training performance remains unanswered, both for the OPF and the general case. We propose a general approach for automated RL environment design by utilizing multi-objective optimization. For that, we use the hyperparameter optimization (HPO) framework, which allows the reuse of existing HPO algorithms and methods. On five OPF benchmark problems, we demonstrate that our automated design approach consistently outperforms a manually created baseline environment design. Further, we use statistical analyses to determine which environment design decisions are especially important for performance, resulting in multiple novel insights on how RL-OPF environments should be designed. Finally, we discuss the risk of overfitting the environment to the utilized RL algorithm. To the best of our knowledge, this is the first general approach for automated RL environment design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。