arXiv:2509.10163cs.LGcs.IT2025-09被引 3

6G边缘网络中用联邦多智能体强化学习实现节能、隐私保护的资源管理。

Federated Multi-Agent Reinforcement Learning for Privacy-Preserving and Energy-Aware Resource Management in 6G Edge Networks

  • 各设备智能体通过局部观察自主决策任务卸载、频谱接入和能耗调节。
  • 相比集中式方法,任务成功率提升18%,延迟降低23%,能耗下降27%。
  • 适合关注6G隐私安全与能效优化的研究者或工程团队。

随着第六代(6G)网络向超密集、智能化边缘环境演进,如何在严格的隐私、移动性和能源约束下实现高效资源管理成为关键挑战。本文提出一种新型联邦多智能体强化学习(Fed-MARL)框架,实现媒体访问控制(MAC)层与应用层的跨层协同,支持异构边缘设备上节能、隐私保护且实时的资源管理。每个智能体采用深度循环Q网络(DRQN)基于本地观测(如队列长度、能量、CPU使用率和移动性)学习去中心化的任务卸载、频谱接入与CPU能耗调节策略。为保障隐私,引入基于椭圆曲线迪菲-赫尔曼密钥交换的安全聚合协议,确保模型更新准确且不泄露原始数据,抵御半诚实攻击。将资源管理问题建模为部分可观测多智能体马尔可夫决策过程(POMMDP),设计多目标奖励函数联合优化延迟、能效、频谱效率、公平性和可靠性,满足6G典型服务需求(URLLC、eMBB、mMTC)。仿真结果表明,与集中式MARL及启发式基线相比,Fed-MARL在任务成功率、延迟、能效和公平性方面均显著提升,同时在动态、资源受限的6G边缘网络中具备强隐私保护能力和可扩展性。

原文摘要 · Abstract (English)

As sixth-generation (6G) networks move toward ultra-dense, intelligent edge environments, efficient resource management under stringent privacy, mobility, and energy constraints becomes critical. This paper introduces a novel Federated Multi-Agent Reinforcement Learning (Fed-MARL) framework that incorporates cross-layer orchestration of both the MAC layer and application layer for energy-efficient, privacy-preserving, and real-time resource management across heterogeneous edge devices. Each agent uses a Deep Recurrent Q-Network (DRQN) to learn decentralized policies for task offloading, spectrum access, and CPU energy adaptation based on local observations (e.g., queue length, energy, CPU usage, and mobility). To protect privacy, we introduce a secure aggregation protocol based on elliptic curve Diffie Hellman key exchange, which ensures accurate model updates without exposing raw data to semi-honest adversaries. We formulate the resource management problem as a partially observable multi-agent Markov decision process (POMMDP) with a multi-objective reward function that jointly optimizes latency, energy efficiency, spectral efficiency, fairness, and reliability under 6G-specific service requirements such as URLLC, eMBB, and mMTC. Simulation results demonstrate that Fed-MARL outperforms centralized MARL and heuristic baselines in task success rate, latency, energy efficiency, and fairness, while ensuring robust privacy protection and scalability in dynamic, resource-constrained 6G edge networks.

6G联邦学习强化学习边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。