用目标空间规划提升强化学习的长期调度能力,解决电力需求响应中的终端约束问题。
Addressing Terminal Constraints in Data-Driven Demand Response Scheduling
- 引入目标空间规划与深度确定性策略梯度结合,通过离散子目标抽象建模实现长程价值传播。
- 在空气分离模拟场景中,样本效率优于标准DDPG,且100%满足终端储气约束。
- 适合关注工业负荷灵活调度、需长期规划与稳定性的能源系统研究者。
电气化化工过程为应对动态电价市场而需灵活运行,但参与需求响应计划常需在长周期内满足终端约束,以保障系统动态稳定性。传统模型优化方法计算成本高,基于强化学习的数据驱动调度面临严重的信用分配难题。本文将目标空间规划(GSP)与深度确定性策略梯度(DDPG)结合,利用对离散子目标的时序抽象模型,实现跨长时间跨度的价值传播。在模拟的空气分离基准任务中,所提方法在保持终端储气约束满足的前提下,显著提升样本效率,缓解了强化学习固有的短视控制行为。
原文摘要 · Abstract (English)
Electrified chemical processes are incentivized by exposure to time-varying electricity markets to operate flexibly, but participating in demand response schemes can require satisfying terminal constraints over long horizons. Specifically, terminal constraints may be required when computing optimal schedules in order to preserve dynamic stability. Model-based optimization methods are computationally costly, and data-driven scheduling via reinforcement learning (RL) faces severe credit-assignment challenges. We integrate Goal-Space Planning (GSP) with Deep Deterministic Policy Gradient (DDPG), using learned temporally abstract models over discrete subgoals to propagate value across extended horizons. Using a simulated air separation benchmark, we demonstrate the proposed approach improves sample efficiency over standard DDPG while satisfying terminal storage constraints, mitigating myopic control behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。