用深度强化学习设计激励机制,提升工业物联网联邦学习的效率与参与度。
Meta-Computing Enhanced Federated Learning in IIoT: Satisfaction-Aware Incentive Scheme via DRL-Based Stackelberg Game
- 基于斯塔克尔伯格博弈建模服务器与节点互动,用深度强化学习求解均衡
- 在相同预算下,系统效用提升至少23.7%,模型精度不下降
- 融合数据量、信息时效性与训练延迟,实现满意度驱动的激励
工业互联网(IIoT)利用联邦学习(FL)进行分布式模型训练以保护数据隐私,元计算通过优化和整合分布式计算资源提升其效率与可扩展性。高效运行需在模型质量与训练延迟间取得平衡。本文设计一个考虑数据大小、信息时效性(AoI)与训练延迟的满意度函数,并将其融入效用函数以激励节点参与。将服务器与节点的效用建模为两阶段斯塔克尔伯格博弈,采用深度强化学习方法学习均衡解,确保奖励均衡且激励方案更具适用性。仿真结果表明,在相同预算约束下,所提方案相较现有无激励的联邦学习方案,系统效用提升至少23.7%,且不牺牲模型准确率。
原文摘要 · Abstract (English)
The Industrial Internet of Things (IIoT) leverages Federated Learning (FL) for distributed model training while preserving data privacy, and meta-computing enhances FL by optimizing and integrating distributed computing resources, improving efficiency and scalability. Efficient IIoT operations require a trade-off between model quality and training latency. Consequently, a primary challenge of FL in IIoT is to optimize overall system performance by balancing model quality and training latency. This paper designs a satisfaction function that accounts for data size, Age of Information (AoI), and training latency for meta-computing. Additionally, the satisfaction function is incorporated into the utility function to incentivize IIoT nodes to participate in model training. We model the utility functions of servers and nodes as a two-stage Stackelberg game and employ a deep reinforcement learning approach to learn the Stackelberg equilibrium. This approach ensures balanced rewards and enhances the applicability of the incentive scheme for IIoT. Simulation results demonstrate that, under the same budget constraints, the proposed incentive scheme improves utility by at least 23.7% compared to existing FL schemes without compromising model accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。