arXiv:2607.12763cs.LGcs.AI2026-07

为微电网能源协调设计安全约束感知的联邦强化学习聚合方法

Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination

论文配图:Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination
图 1 · 摘自论文原文
  • 在服务器端融合局部奖励与约束违反估计,设计轻量级聚合规则
  • 惩罚机制使奖励提升12%且违规次数减少67%,优于传统FedAvg
  • 无需修改本地训练,适合实际部署于分布式能源系统

联邦强化学习(FedRL)可在不共享原始数据的前提下协调分布式能源资源,但标准聚合方法如FedAvg忽视系统级约束,常导致不安全的全局行为。本文研究微电网能源协调中的约束感知聚合,提出将局部性能与约束违反估计结合的服务器端更新规则。其中,基于惩罚的简单规则 $w_i /propto R_i - αV_i$ 在奖励与安全性之间实现最稳定的权衡,无需对偶优化或修改本地训练。我们在DairyGridEnv基准上评估该方法,该基准模拟多个农场在随机需求和共享电网容量约束下协调储能;进一步使用芬兰和德国FIELD数据集的真实负荷驱动需求进行鲁棒性测试。在多个随机种子下,惩罚聚合显著减少违规次数,同时相比FedAvg提升奖励表现。联合奖赏-违规方案通过$λ$调节权衡,但稳定性较差。结果表明,轻量级聚合策略可大幅提升联邦强化学习的实证安全性,同时保持标准通信协议。

原文摘要 · Abstract (English)

Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation methods such as FedAvg do not account for system-level constraints, often leading to unsafe global behavior. In this work, we study constraint-aware aggregation for federated reinforcement learning in distributed energy coordination. We propose aggregation rules that incorporate both local performance and estimated constraint violation into the server-side update. Among these, a simple penalty-based rule, $w_i \propto R_i - αV_i$, consistently provides the most reliable trade-off between reward and safety, without requiring dual optimization or modifications to local training. \textcolor{black}{We evaluate our approach on DairyGridEnv, a benchmark modeling multiple farms coordinating battery storage under stochastic demand and a shared grid capacity constraint, and further assess robustness using real load-driven demand profiles from Finland and the German FIELD dataset. Across multiple seeds, penalty-based aggregation substantially reduces violations while improving reward relative to FedAvg in both synthetic and real load-driven settings.} A combined reward-violation scheme exposes a tunable trade-off via $λ$, but is less stable. These results demonstrate that lightweight aggregation strategies can substantially improve empirical safety in federated reinforcement learning while preserving standard communication protocols.

联邦学习强化学习能源系统安全约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。