arXiv:2511.02646cs.LGcs.AI2025-11

用深度强化学习模拟天然气存储,预测价格波动并评估政策效果。

Natural-gas storage modelling by deep reinforcement learning

  • 用深度强化学习训练储气运营商策略,耦合真实市场数据。
  • SAC算法使储气目标、价格稳定等多重目标同时达成,匹配真实价格波动特征。
  • 无需直接拟合价格数据,即可还原历史价格分布,适合能源政策评估。

我们提出GasRL,一个将校准的天然气市场模型与基于深度强化学习(RL)训练的储气运营策略相结合的仿真器。通过该模型分析最优储气管理对均衡价格及供需动态的影响。测试多种RL算法后发现,软演员-评论家(SAC)在该环境中表现最优:成功实现储气运营商的多重目标——包括盈利性、市场清算稳健性与价格稳定性。此外,由SAC生成的最优策略所引致的均衡价格动态,具备与真实市场价格高度吻合的波动性与季节性特征。值得注意的是,这一对历史价格分布的拟合,是在未显式校准至价格数据的前提下实现的。我们还展示了该仿真器在评估欧盟强制最低储气阈值政策效果中的应用。结果表明,此类阈值能提升市场对意外供应冲击分布变化的韧性;例如,在遭遇异常大的冲击时,有阈值的市场更常避免发生中断。

原文摘要 · Abstract (English)

We introduce GasRL, a simulator that couples a calibrated representation of the natural gas market with a model of storage-operator policies trained with deep reinforcement learning (RL). We use it to analyse how optimal stockpile management affects equilibrium prices and the dynamics of demand and supply. We test various RL algorithms and find that Soft Actor Critic (SAC) exhibits superior performance in the GasRL environment: multiple objectives of storage operators - including profitability, robust market clearing and price stabilisation - are successfully achieved. Moreover, the equilibrium price dynamics induced by SAC-derived optimal policies have characteristics, such as volatility and seasonality, that closely match those of real-world prices. Remarkably, this adherence to the historical distribution of prices is obtained without explicitly calibrating the model to price data. We show how the simulator can be used to assess the effects of EU-mandated minimum storage thresholds. We find that such thresholds have a positive effect on market resilience against unanticipated shifts in the distribution of supply shocks. For example, with unusually large shocks, market disruptions are averted more often if a threshold is in place.

强化学习能源市场储气模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。