arXiv:2512.01046cs.AIcs.LG2025-12

用可解释的防护单元让强化学习在微电网中安全运行

Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids

  • 将系统约束拆解为层级化防护单元,结合先验知识确保合规
  • 实测燃料消耗降24%且电池损耗不增,满足所有运行限制
  • 适合需高安全性的能源系统决策场景

强化学习(RL)是应对复杂系统中不确定性决策的强大框架,尤其在能源转型背景下至关重要。以远离主网的偏远微电网为例,需协调风电、柴油发电机与电池储能,在外部间歇性负荷与风力条件下满足用电需求,同时最小化燃油消耗和电池退化。此类系统常受严格监管和复杂运行约束制约。为确保RL代理遵守这些约束,提供可解释的保障至关重要。本文提出屏蔽控制器单元(SCUs),一种系统化且可解释的方法,利用系统动态先验知识保证约束满足。其屏蔽合成方法专为实际部署设计,将环境分解为分层结构,每个SCU显式管理一组约束。我们在具有严格运行要求的微电网优化任务上验证了SCUs的有效性:配备SCUs的RL代理实现24%燃油消耗降低,且未增加电池退化,优于其他基线方法,同时满足所有约束。我们期望SCUs能推动强化学习在能源转型相关决策挑战中的安全应用。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is a powerful framework for optimizing decision-making in complex systems under uncertainty, an essential challenge in real-world settings, particularly in the context of the energy transition. A representative example is remote microgrids that supply power to communities disconnected from the main grid. Enabling the energy transition in such systems requires coordinated control of renewable sources like wind turbines, alongside fuel generators and batteries, to meet demand while minimizing fuel consumption and battery degradation under exogenous and intermittent load and wind conditions. These systems must often conform to extensive regulations and complex operational constraints. To ensure that RL agents respect these constraints, it is crucial to provide interpretable guarantees. In this paper, we introduce Shielded Controller Units (SCUs), a systematic and interpretable approach that leverages prior knowledge of system dynamics to ensure constraint satisfaction. Our shield synthesis methodology, designed for real-world deployment, decomposes the environment into a hierarchical structure where each SCU explicitly manages a subset of constraints. We demonstrate the effectiveness of SCUs on a remote microgrid optimization task with strict operational requirements. The RL agent, equipped with SCUs, achieves a 24% reduction in fuel consumption without increasing battery degradation, outperforming other baselines while satisfying all constraints. We hope SCUs contribute to the safe application of RL to the many decision-making challenges linked to the energy transition.

强化学习微电网安全控制能源转型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。