arXiv:2512.02350cs.LGcs.AI2025-12中稿 · IEEE/ACM ToN被引 4

解决联邦强化学习中低质量数据干扰问题,提升算法稳定性与性能

FOVA: Offline Federated Reinforcement Learning with Mixed-Quality Data

  • 引入投票机制筛选高回报动作,缓解不同客户端数据质量差异影响
  • 基于优势加权回归构建一致训练目标,显著提升训练效率与稳定性
  • 理论证明策略严格优于行为策略,适合多源异构数据场景

离线联邦强化学习(FRL)结合了联邦学习与离线强化学习的优势,近年来受到广泛关注。然而,现有方法在面对混合质量数据时表现显著下降——即各客户端的离线数据由不同质量的策略生成。为此,本文提出一种基于投票机制的离线联邦强化学习框架FOVA。它通过投票机制识别局部策略评估中的高回报动作,有效缓解低质量行为带来的负面影响。同时,基于优势加权回归(AWR),构建了一致的局部与全局训练目标,大幅提升训练效率与稳定性。进一步的理论分析严格证明,FOVA学习到的策略严格优于行为策略。大量实验表明,该算法在主流基准上显著优于现有基线。

原文摘要 · Abstract (English)

Offline Federated Reinforcement Learning (FRL), a marriage of federated learning and offline reinforcement learning, has attracted increasing interest recently. Albeit with some advancement, we find that the performance of most existing offline FRL methods drops dramatically when provided with mixed-quality data, that is, the logging behaviors (offline data) are collected by policies with varying qualities across clients. To overcome this limitation, this paper introduces a new vote-based offline FRL framework, named FOVA. It exploits a \emph{vote mechanism} to identify high-return actions during local policy evaluation, alleviating the negative effect of low-quality behaviors from diverse local learning policies. Besides, building on advantage-weighted regression (AWR), we construct consistent local and global training objectives, significantly enhancing the efficiency and stability of FOVA. Further, we conduct an extensive theoretical analysis and rigorously show that the policy learned by FOVA enjoys strict policy improvement over the behavioral policy. Extensive experiments corroborate the significant performance gains of our proposed algorithm over existing baselines on widely used benchmarks.

联邦学习强化学习离线学习数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。