提出可微分的双正性参数化方法,提升电力市场强化学习模拟的可信度
A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

- 设计双正性参数化实现单调有界多段报价的可微生成
- 在3个真实市场数据集上验证模拟结果接近纳什均衡
- 为强化学习电力市场仿真提供严谨的可行性评估框架
强化学习代理仿真(RL-ABS)已成为电力市场机制分析与评估的重要工具。在建模单调、有界、多段阶梯式报价时,现有方法通常先让策略网络输出无约束动作,再通过排序、截断或投影等后处理映射转换为满足单调性和有界性的报价曲线。然而,这类后处理映射在边界或拐点处往往不满足连续可微性、单射性和可逆性,导致梯度畸变,引发模拟结果的虚假收敛。同时,多数现有研究仅依据训练曲线收敛进行机制分析与评估,未严格衡量仿真结果与纳什均衡之间的距离,严重削弱了结论可信度。为此,本文提出一种双正性单调参数化方法,确保报价曲线生成过程全程可微且可逆,并构建一个基于纳什均衡距离的可行性评估框架。实验在3个真实电力市场数据集上验证,新方法显著提升仿真结果的合理性与收敛稳定性,为电力市场强化学习仿真提供更可靠的分析基础。
原文摘要 · Abstract (English)
Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the modeling of monotone, bounded, multi-segment stepwise bids, existing methods typically let the policy network first output an unconstrained action and then convert it into a feasible bid curve satisfying monotonicity and boundedness through post-processing mappings such as sorting, clipping, or projection. However, such post-processing mappings often fail to satisfy continuous differentiability, injectivity, and invertibility at boundaries or kinks, thereby causing gradient distortion and leading to spurious convergence in simulation results. Meanwhile, most existing studies conduct mechanism analysis and evaluation mainly on the basis of training-curve convergence, without rigorously assessing the distance between the simulation outcomes and Nash equilibrium, which severely undermines the credibility of the results. To address these issues, this paper proposes...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。