用强化学习优化电池群调度,降低衰减并提升电网响应能力
Degradation-Aware Frequency Regulation of a Heterogeneous Battery Fleet via Reinforcement Learning
- 构建马尔可夫决策模型,通过代理奖励反馈实时调控电池充放电
- 相比基线策略,循环深度下降超30%,寿命衰减显著减少
- 适合电力系统、储能管理与智能调度方向的研究者参考
电池储能系统被广泛用于电网频率调节及缓解可再生能源波动。然而频繁充放电会引发循环衰减,缩短电池寿命。本文研究异构电池群在随机调节信号下的实时调度问题,在每节电池的爬坡率与容量约束下,最小化长期循环衰减。由于衰减具有路径依赖性(由荷电状态轨迹决定),传统动态规划难以建模。为此,将问题建模为带约束动作空间的马尔可夫决策过程,并设计密集代理奖励,在每步提供有效反馈的同时优化长期循环深度。为应对精细荷电状态离散化和不对称约束带来的高维状态-动作空间,采用极限学习机作为随机非线性特征映射,结合线性时序差分学习进行函数逼近。在模拟马尔可夫信号及从特拉华大学实测调节信号训练得到的马尔可夫模型上评估,结果表明该方法持续降低循环深度出现频率与衰减指标,优于基准调度策略。
原文摘要 · Abstract (English)
Battery energy storage systems are increasingly deployed as fast-responding resources for grid balancing services such as frequency regulation and for mitigating renewable generation uncertainty. However, repeated charging and discharging induces cycling degradation and reduces battery lifetime. This paper studies the real-time scheduling of a heterogeneous battery fleet that collectively tracks a stochastic balancing signal subject to per-battery ramp-rate and capacity constraints, while minimizing long-term cycling degradation. Cycling degradation is fundamentally path-dependent: it is determined by charge-discharge cycles formed by the state-of-charge (SoC) trajectory and is commonly quantified via rainflow cycle counting. This non-Markovian structure makes it difficult to express degradation as an additive per-time-step cost, complicating classical dynamic programming approaches. We address this challenge by formulating the fleet scheduling problem as a Markov decision process (MDP) with constrained action space and designing a dense proxy reward that provides informative feedback at each time step while remaining aligned with long-term cycle-depth reduction. To scale learning to large state-action spaces induced by fine-grained SoC discretization and asymmetric per-battery constraints, we develop a function-approximation reinforcement learning method using an Extreme Learning Machine (ELM) as a random nonlinear feature map combined with linear temporal-difference learning. We evaluate the proposed approach on a toy Markovian signal model and on a Markovian model trained from real-world regulation signal traces obtained from the University of Delaware, and demonstrate consistent reductions in cycle-depth occurrence and degradation metrics compared to baseline scheduling policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。