arXiv:2409.05291cs.LGcs.SY2024-09中稿 · the Decision and C…被引 3

提出Fast-FedPG算法,实现联邦强化学习的快速收敛与无偏差优化。

Towards Fast Rates for Federated and Multi-Task Reinforcement Learning

  • 设计带偏差校正的联邦策略梯度算法,解决异构任务下的协作难题。
  • 在理想条件下线性收敛,在噪声条件下实现随代理数线性加速的次线性收敛。
  • 适用于多智能体协同强化学习场景,尤其适合目标异构的分布式系统。

我们研究一个包含N个智能体的设置,每个智能体与一个马尔可夫决策过程(MDP)交互,其奖励函数不同,体现异构的目标或任务。各智能体通过中心服务器间歇通信,共同寻找最大化各环境长期累积奖励平均值的策略。现有工作要么仅提供渐近速率,要么产生有偏策略,或未能证明协作优势。为此,我们提出Fast-FedPG——一种具有精心设计偏差校正机制的新型联邦策略梯度算法。在梯度支配条件下,我们证明该算法在使用精确梯度时实现快速线性收敛,并在使用带有噪声和截断的策略梯度时获得随智能体数线性加速的次线性速率。值得注意的是,两种情况下收敛均达到全局最优策略且无异构引入的偏差。在缺乏梯度支配条件时,仍可保证以持续受益于协作的速率收敛至一阶驻点。

原文摘要 · Abstract (English)

We consider a setting involving $N$ agents, where each agent interacts with an environment modeled as a Markov Decision Process (MDP). The agents' MDPs differ in their reward functions, capturing heterogeneous objectives/tasks. The collective goal of the agents is to communicate intermittently via a central server to find a policy that maximizes the average of long-term cumulative rewards across environments. The limited existing work on this topic either only provide asymptotic rates, or generate biased policies, or fail to establish any benefits of collaboration. In response, we propose Fast-FedPG - a novel federated policy gradient algorithm with a carefully designed bias-correction mechanism. Under a gradient-domination condition, we prove that our algorithm guarantees (i) fast linear convergence with exact gradients, and (ii) sub-linear rates that enjoy a linear speedup w.r.t. the number of agents with noisy, truncated policy gradients. Notably, in each case, the convergence is to a globally optimal policy with no heterogeneity-induced bias. In the absence of gradient-domination, we establish convergence to a first-order stationary point at a rate that continues to benefit from collaboration.

联邦学习强化学习策略梯度多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。