arXiv:2507.07320cs.LG2025-07被引 17

提升隐私保护下集群联邦学习的通信效率与训练效果

Optimizing Communication and Device Clustering for Clustered Federated Learning with Differential Privacy

  • 提出动态惩罚机制的多智能体强化学习算法,优化设备分组与资源分配
  • 相比独立Q-learning,收敛速度提升20%,累积奖励提高15%
  • 适用于异构基站、非独立同分布数据场景,兼顾隐私与通信约束

本文提出一种安全且通信高效的集群联邦学习(CFL)设计。在该模型中,具备异构任务处理能力的多个基站(BSs)与具有非独立同分布(non-IID)数据的多个用户共同执行融合差分隐私(DP)技术的CFL训练。由于每个基站只能处理部分学习任务,且无线资源块(RBs)有限,需联合优化RB分配与用户调度以提升CFL性能。同时,设备需利用有限数据和模型信息判断任务身份,可能引入额外通信开销。我们构建优化问题,目标是最小化所有任务的训练损失,同时考虑设备聚类、RB分配、DP噪声及模型传输延迟。为此,提出一种新型动态惩罚函数辅助的值分解多智能体强化学习(DPVD-MARL)算法,使各基站可独立决定连接用户、资源块及噪声,但全局最小化所有任务的训练损失。不同于传统方法对无效动作施加固定大惩罚,本方案根据无法满足通信约束(如延迟)的设备数量动态分配惩罚,加速找到有效策略,提升收敛速度。仿真结果表明,相比独立Q-learning,DPVD-MARL可提升收敛率最高达20%,最终累积奖励提高15%。

原文摘要 · Abstract (English)

In this paper, a secure and communication-efficient clustered federated learning (CFL) design is proposed. In our model, several base stations (BSs) with heterogeneous task-handling capabilities and multiple users with non-independent and identically distributed (non-IID) data jointly perform CFL training incorporating differential privacy (DP) techniques. Since each BS can process only a subset of the learning tasks and has limited wireless resource blocks (RBs) to allocate to users for federated learning (FL) model parameter transmission, it is necessary to jointly optimize RB allocation and user scheduling for CFL performance optimization. Meanwhile, our considered CFL method requires devices to use their limited data and FL model information to determine their task identities, which may introduce additional communication overhead. We formulate an optimization problem whose goal is to minimize the training loss of all learning tasks while considering device clustering, RB allocation, DP noise, and FL model transmission delay. To solve the problem, we propose a novel dynamic penalty function assisted value decomposed multi-agent reinforcement learning (DPVD-MARL) algorithm that enables distributed BSs to independently determine their connected users, RBs, and DP noise of the connected users but jointly minimize the training loss of all learning tasks across all BSs. Different from the existing MARL methods that assign a large penalty for invalid actions, we propose a novel penalty assignment scheme that assigns penalty depending on the number of devices that cannot meet communication constraints (e.g., delay), which can guide the MARL scheme to quickly find valid actions, thus improving the convergence speed. Simulation results show that the DPVD-MARL can improve the convergence rate by up to 20% and the ultimate accumulated rewards by 15% compared to independent Q-learning.

联邦学习差分隐私多智能体通信优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。