arXiv:2412.15538cs.LGcs.AI2024-12被引 18

FedRLHF让多方协作训练强化学习模型,不共享数据也能保证隐私和个性化。

FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF

  • 各客户端本地整合人类反馈优化奖励函数,通过联邦学习协同更新策略。
  • 理论证明算法收敛,样本复杂度随客户端数量高效增长。
  • 在电影和影评数据集上表现媲美中心化方法,适合隐私敏感场景。

随着隐私关注日益增加与个性化需求上升,传统基于人类反馈的强化学习(RLHF)框架因依赖集中式数据而面临严峻挑战。本文提出联邦人类反馈强化学习(FedRLHF),一种去中心化的新型框架,使多个客户端可在不共享原始数据或人类反馈的前提下协作进行策略学习,从而实现强隐私保护。该框架利用联邦强化学习技术,使每个客户端在本地将人类反馈融入奖励函数,并通过个性化RLHF过程更新策略。我们为FedRLHF建立了严格的理论基础,提供了收敛性保证,并推导出随客户端数量高效增长的样本复杂度边界。在MovieLens和IMDb数据集上的实证评估表明,FedRLHF不仅有效保护用户隐私,还能达到与集中式RLHF相当的性能,同时在多样化客户端环境中显著提升个性化水平。

原文摘要 · Abstract (English)

In the era of increasing privacy concerns and demand for personalized experiences, traditional Reinforcement Learning with Human Feedback (RLHF) frameworks face significant challenges due to their reliance on centralized data. We introduce Federated Reinforcement Learning with Human Feedback (FedRLHF), a novel framework that decentralizes the RLHF process. FedRLHF enables collaborative policy learning across multiple clients without necessitating the sharing of raw data or human feedback, thereby ensuring robust privacy preservation. Leveraging federated reinforcement learning, each client integrates human feedback locally into their reward functions and updates their policies through personalized RLHF processes. We establish rigorous theoretical foundations for FedRLHF, providing convergence guarantees, and deriving sample complexity bounds that scale efficiently with the number of clients. Empirical evaluations on the MovieLens and IMDb datasets demonstrate that FedRLHF not only preserves user privacy but also achieves performance on par with centralized RLHF, while enhancing personalization across diverse client environments.

联邦学习强化学习隐私保护个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。