用内在动机驱动探索,提升联邦强化学习的个性化与效率。
Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

- 客户端通过内在好奇心机制主动探索未知状态空间。
- 在稀疏奖励环境下,策略个性化与样本效率显著优于基准方法。
- 保护隐私:服务器仅接收最小新颖性摘要,不接触原始数据。
个性化联邦强化学习(PFRL)在保持客户端数据隐私的前提下,基于历史经验进行分布式学习。现有方法过度依赖外在奖励信号,忽视非平稳或稀疏奖励环境中的探索。本文提出基于内在动机的探索驱动框架EDPFRL-IM,通过引入内在随机网络蒸馏(RND)信号增强客户端本地探索能力。服务器不获取客户端原始经验或局部梯度,仅发送全局探索先验,并收集各客户端的最小新颖性摘要,以实现跨客户端的多样且协调的探索。在基准环境中的实验表明,该框架在延迟与稀疏奖励系统中显著优于平均PFRL基线,在策略个性化和样本效率方面表现更优。整体上,EDPFRL-IM实现了可灵活扩展的探索学习结构与客户端隐私保护的融合。
原文摘要 · Abstract (English)
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments. In this work, we introduce a new exploration-driven framework, Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), that leverages an inherent curiosity-driven exploration at each client to promote local exploration and protect client privacy. Furthermore, to facilitate policy discovery via exploration in previously unexplored state spaces, clients add an intrinsic random network distillation (RND) signal to their extrinsic reward. Additionally, the server does not have access to clients' raw experiences or local gradient estimates; instead, the server sends global exploration priors and collects minimal novelty summaries from each client to enable both diverse and coordinated exploration among clients. Experiments in benchmark environments show that our framework outperforms average PFRL benchmarks in policy personalization and sample efficiency, primarily in delayed and sparse reward systems. Overall, EDPFRL-IM enables the integration of a flexible exploratory learning structure into federated reinforcement learning systems while preserving client privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。