用公共数据提升联邦强化推理的协作效率与一致性。
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR

- 基于LoRA的本地适配+公共数据离策略更新,降低通信开销。
- 在数学与医疗推理任务中,显著优于传统基线方法。
- 适合隐私敏感场景下跨组织模型协同优化,如医疗、金融。
基于可验证奖励的强化学习推理(RLVR)通常在集中式环境中研究,但现实应用多涉及分布于各组织的私有数据。联邦训练是自然解决方案,但在此范式中扩展RLVR面临挑战:全模型同步成本高,大量本地迭代易导致客户端漂移。本文提出一种联邦RLVR框架,结合基于LoRA的本地适应与基于公共数据的离策略步骤,提升通信效率与跨客户端协调能力。具体而言,使用小型共享公共数据集周期性交换和复用响应级训练信号,提供轻量级全局对齐锚点,且不暴露私有数据。在公共数据步中,选择性替换本地错误响应为全局正确响应,保持训练贴近本地策略的同时受益于跨客户端协作。在数学与医疗推理基准及模型上,本方法持续优于标准基线。结果表明,低秩通信与有限公共数据协调的组合是一种简单而有效的联邦推理后训练方案。
原文摘要 · Abstract (English)
Reasoning post-training with reinforcement learning from verifiable rewards (RLVR) is typically studied in centralized settings, yet many realistic applications involve decentralized private data distributed across organizations. Federated training is a natural solution, but scaling RLVR in this regime is challenging: full-model synchronization is expensive, and performing many local steps can cause severe client drift under heterogeneous data. We propose a federated RLVR framework that combines LoRA-based local adaptation with public-data-based off-policy steps to improve both communication efficiency and cross-client coordination. In particular, a small shared public dataset is used to periodically exchange and reuse response-level training signals across organizations, providing a lightweight anchor toward a more globally aligned objective without exposing private data. Our method selectively replaces locally incorrect responses with globally correct ones during public-data steps, thereby keeping training closer to the local policy while still benefiting from cross-client coordination. Across mathematical and medical reasoning benchmarks and models, our method consistently improves over standard baselines. Our results highlight a simple and effective recipe for federated reasoning post-training: combining low-rank communication with limited public-data coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。