面向边缘设备的高效联邦强化学习反馈优化方法
Efficient Federated RLHF via Zeroth-Order Policy Optimization
- 采用零阶优化与二值扰动,降低通信计算开销
- 在四个MuJoCo任务中优于基线方法,收敛更快
- 适合资源受限的分布式智能系统应用
本文研究在资源受限的边缘设备上进行联邦强化学习从人类反馈(RLHF)的问题。提出一种名为分区、基于符号的随机零阶策略优化(Par-S²ZPO)的高效联邦RLHF算法。该算法基于零阶优化与二值扰动设计,天然具备低通信、低计算和低内存复杂度。理论分析表明,其收敛速率上界显示,在样本复杂度上与集中式方法相当,但在策略更新迭代次数上收敛更快。实验结果表明,该算法在四个MuJoCo强化学习任务上均优于基于FedAvg的RLHF基线方法。
原文摘要 · Abstract (English)
This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices. We propose an efficient federated RLHF algorithm, named Partitioned, Sign-based Stochastic Zeroth-order Policy Optimization (Par-S$^2$ZPO). The algorithm is built on zeroth-order optimization with binary perturbation, resulting in low communication, computation, and memory complexity by design. Our theoretical analysis establishes an upper bound on the convergence rate of Par-S$^2$ZPO, revealing that it is as efficient as its centralized counterpart in terms of sample complexity but converges faster in terms of policy update iterations. Our experimental results show that it outperforms a FedAvg-based RLHF on four MuJoCo RL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。