arXiv:2508.16037cs.LGcs.AI2025-08被引 4

多服务提供商联邦学习中,用强化学习协调通信与计算资源,实现各方利益平衡。

Pareto Actor-Critic for Communication and Computation Co-Optimization in Non-Cooperative Federated Learning Services

  • 基于博弈论的多智能体强化学习框架,联合优化任务分配与资源调度。
  • 相比最新方法,总奖励提升5.8%,系统性能指标提高4.2%。
  • 适合大规模异构数据环境下的非合作联邦学习系统设计与部署。

在多服务提供商(SP)生态系统的联邦学习中,隐私限制与利益竞争导致通信与计算资源无法集中优化。本文提出PAC-MCoFL框架,将各SP建模为智能体,通过集成帕累托-演员评论家(PAC)与期望回归机制,在建模异质风险偏好的基础上,协同优化客户端分配、自适应量化与资源分配。为应对高维动作空间,设计三元笛卡尔分解(TCAD)机制实现细粒度控制。进一步提出可扩展的PAC-MCoFL-p版本,采用参数化假设生成器显著降低计算复杂度,且误差有理论保证。大量仿真验证表明,该框架在总奖励和超体积指标(HVI)上分别优于最新MARL方案约5.8%和4.2%。结果还显示其在规模化部署及不同数据异构性下,能更有效平衡单个SP与整体系统性能。

原文摘要 · Abstract (English)

Federated learning (FL) in multi-service provider (SP) ecosystems is fundamentally hampered by non-cooperative dynamics, where privacy constraints and competing interests preclude the centralized optimization of multi-SP communication and computation resources. In this paper, we introduce PAC-MCoFL, a game-theoretic multi-agent reinforcement learning (MARL) framework where SPs act as agents to jointly optimize client assignment, adaptive quantization, and resource allocation. Within the framework, we integrate Pareto Actor-Critic (PAC) principles with expectile regression, enabling agents to conjecture optimal joint policies to achieve Pareto-optimal equilibria while modeling heterogeneous risk profiles. To manage the high-dimensional action space, we devise a ternary Cartesian decomposition (TCAD) mechanism that facilitates fine-grained control. Further, we develop PAC-MCoFL-p, a scalable variant featuring a parameterized conjecture generator that substantially reduces computational complexity with a provably bounded error. Alongside theoretical convergence guarantees, our framework's superiority is validated through extensive simulations -- PAC-MCoFL achieves approximately 5.8% and 4.2% improvements in total reward and hypervolume indicator (HVI), respectively, over the latest MARL solutions. The results also demonstrate that our method can more effectively balance individual SP and system performance in scaled deployments and under diverse data heterogeneity.

联邦学习强化学习资源优化多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。