arXiv:2503.17454cs.LG2025-03被引 1

联邦强化学习在环境不一致时仍能合作估计价值函数。

Collaborative Value Function Estimation Under Model Mismatch: A Federated Temporal Difference Analysis

  • 分析单智能体和联邦版本的时序差分学习收敛性。
  • 模型偏差导致估值不准,但共享信息可显著降低误差。
  • 适合关注隐私保护下多智能体协作的研究者。

联邦强化学习(FedRL)通过避免直接交换数据,在保护隐私的前提下实现协作学习。然而,现有算法通常假设所有智能体处于相同环境中,这在现实场景中往往不成立。例如在多机器人团队、众包系统和大规模传感器网络中,各智能体可能面临略有不同的转移动态,导致固有的模型不匹配。本文首先建立单智能体时序差分学习(TD(0))在策略评估中的线性收敛性,并证明在扰动环境下,智能体会产生系统性偏差,无法准确估计真实价值函数,该结论在独立同分布与马尔可夫采样情形下均成立。随后将分析扩展至联邦TD(0)(FedTD(0))设置,多个智能体各自与扰动环境交互,定期共享价值估计以协同逼近共同底层模型的真实价值函数。理论结果揭示了模型不匹配、网络连通性与混合行为对FedTD(0)收敛的影响。实验验证了理论发现,表明即使适度的信息共享也能显著缓解环境特异性误差。

原文摘要 · Abstract (English)

Federated reinforcement learning (FedRL) enables collaborative learning while preserving data privacy by preventing direct data exchange between agents. However, many existing FedRL algorithms assume that all agents operate in identical environments, which is often unrealistic. In real-world applications, such as multi-robot teams, crowdsourced systems, and large-scale sensor networks, each agent may experience slightly different transition dynamics, leading to inherent model mismatches. In this paper, we first establish linear convergence guarantees for single-agent temporal difference learning (TD(0)) in policy evaluation and demonstrate that under a perturbed environment, the agent suffers a systematic bias that prevents accurate estimation of the true value function. This result holds under both i.i.d. and Markovian sampling regimes. We then extend our analysis to the federated TD(0) (FedTD(0)) setting, where multiple agents, each interacting with its own perturbed environment, periodically share value estimates to collaboratively approximate the true value function of a common underlying model. Our theoretical results indicate the impact of model mismatch, network connectivity, and mixing behavior on the convergence of FedTD(0). Empirical experiments corroborate our theoretical gains, highlighting that even moderate levels of information sharing significantly mitigate environment-specific errors.

联邦学习强化学习价值估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。