为联邦学习设计自适应激励机制,提升各方参与积极性。
A Service-Oriented Adaptive Hierarchical Incentive Mechanism for Federated Learning
- 用斯塔克尔伯格博弈建模任务发布方与本地模型方的互动。
- 通过深度强化学习求解工人最优策略,提升数据贡献效率。
- 适合关注联邦学习激励设计与多方协作的研究者。
联邦学习(FL)作为一种分布式训练框架受到广泛关注。在该框架中,任务发布方(TP)发布任务,本地模型所有者(LMOs)利用本地数据训练模型。当存在数据不足时,需招募工人采集数据。为此,本文从服务化视角提出一种自适应激励机制,目标是最大化任务发布方、本地模型所有者和工人的效用。具体地,构建了以任务发布方为领导者、本地模型所有者为追随者的斯塔克尔伯格博弈,并推导出解析的纳什均衡解以优化双方效用。同时,采用多智能体马尔可夫决策过程(MAMDP)建模本地模型所有者与工人之间的交互,通过深度强化学习(DRL)识别最优策略。此外,设计了自适应搜索最优策略算法(ASOSA),用于稳定各参与方策略并解决耦合问题。大量数值实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Recently, federated learning (FL) has emerged as a novel framework for distributed model training. In FL, the task publisher (TP) releases tasks, and local model owners (LMOs) use their local data to train models. Sometimes, FL suffers from the lack of training data, and thus workers are recruited for gathering data. To this end, this paper proposes an adaptive incentive mechanism from a service-oriented perspective, with the objective of maximizing the utilities of TP, LMOs and workers. Specifically, a Stackelberg game is theoretically established between the LMOs and TP, positioning TP as the leader and the LMOs as followers. An analytical Nash equilibrium solution is derived to maximize their utilities. The interaction between LMOs and workers is formulated by a multi-agent Markov decision process (MAMDP), with the optimal strategy identified via deep reinforcement learning (DRL). Additionally, an Adaptively Searching the Optimal Strategy Algorithm (ASOSA) is designed to stabilize the strategies of each participant and solve the coupling problems. Extensive numerical experiments are conducted to validate the efficacy of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。