arXiv:2606.03094cs.LG2026-06

联邦强化学习新框架,隐私保护下实现跨设备高效推理模型训练

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

  • 采用无评估网络的分组相对策略优化,降低计算开销
  • 通过相对收益自适应聚合,解决异构任务间奖励尺度差异问题
  • 适合数据分布不均的分布式场景,尤其关注隐私安全的研究者

语言模型的最新进展确立了强化学习作为激发自我修正与长链推理的主要范式。尽管分组相对策略优化(GRPO)通过消除评价网络实现了更优的可扩展性,但在中心化基础设施上部署仍需汇集大量来自分布式数据持有者的数据,带来显著隐私风险。为此,我们提出联邦GRPO(FGRPO),一种在异构数据持有者之间去中心化微调推理模型的框架。为有效缓解由异构任务导致的奖励尺度差异带来的不稳定性,FGRPO引入基于相对性能提升的自适应聚合机制。通过刻画每个客户端相对于其个性化历史基线的改进程度,该框架可动态优先选择有效学习路径,无论本地任务难度如何。FGRPO在非独立同分布(non-IID)数据上保证了稳健收敛,同时维护数据隐私。

原文摘要 · Abstract (English)

Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative policy optimization (GRPO) offers superior scalability by eliminating the critic network, deploying it on a central infrastructure entails collecting a large volume of data from distributed owners, which poses significant privacy risks. To address these concerns, we introduce federated GRPO (FGRPO), a framework designed to decentralize the fine-tuning of reasoning models across heterogeneous data owners. To effectively mitigate the instability caused by divergent reward scales across heterogeneous tasks, FGRPO incorporates an adaptive aggregation mechanism based on relative performance gain. By characterizing each client's improvement relative to its personalized historical baseline, the framework dynamically prioritizes effective learning trajectories regardless of local task difficulty. FGRPO ensures robust convergence on non-IID data while preserving data privacy.

联邦学习强化学习推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。