用强化学习设计太阳能政策,平衡推广效果与财政支出。
Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

- 将政策制定建模为序贯决策问题,结合强化学习与随机代理模型。
- 最优政策可使4145户采用光伏,成本4173万欧;最低成本仅727万欧,覆盖2682户。
- 算法结果稳定,适合需动态调整的能源政策制定者参考。
设计高效且财政可持续的太阳能光伏发电(PV)推广政策,需在不确定性与异质决策下权衡采纳收益与公共支出。本研究将光伏政策设计建模为序贯决策问题,融合强化学习(RL)与随机代理基础模型(ABM),模拟每年在不确定条件下的光伏采纳行为。一个政策制定者代理在16年期内选择年度激励措施,包括资本补贴、优惠贷款利率和上网电价。通过标量奖励框架调整政策偏好,探索采纳与成本的权衡。采用PPO、SAC和TD3算法学习策略,并在随机模拟中评估。结果表明:最高采纳政策(TD3,$w_{\text{cost}}=0.5$)达约4,145户采纳,成本4173万欧元;最低成本政策(PPO,$w_{\text{cost}}=2.0$)将支出降至727万欧元,覆盖2,682户;均衡政策(PPO,$w_{\text{cost}}=1.6$)实现3,495户采纳,成本2247万欧元。各算法均呈现一致的权衡模式,表明采纳-成本关系稳健。相比静态基线政策,该框架探索了更广泛的政策配置空间。研究证明强化学习在不确定性下具有灵活的自适应政策设计潜力。
原文摘要 · Abstract (English)
Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent selects annual incentives, including capital grants, subsidised loan rates, and feed-in tariffs, over a 16-year horizon. Adoption--cost trade-offs are explored by varying policy preferences within a scalarised reward framework. Policies are learned using PPO, SAC, and TD3 and evaluated under stochastic simulation. The results show that this approach produces a clear trade-off structure: the highest-adoption policy (TD3, $w_{\text{cost}}=0.5$) achieves approximately 4,145 adopters at a cost of EUR 41.73 million, while the lowest-cost policy (PPO, $w_{\text{cost}}=2.0$) reduces expenditure to EUR 7.27 million with 2,682 adopters. The balanced policy (PPO, $w_{\text{cost}}=1.6$) achieves 3,495 adopters at a cost of EUR 22.47 million. Across algorithms, consistent trade-off patterns are observed, indicating robustness of the adoption--cost relationship. Compared with static baseline policies, the RL framework explores a broader range of policy configurations. These findings demonstrate the potential of RL as a flexible tool for adaptive policy design under uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。