arXiv:2409.18718cs.NIcs.LG2024-09被引 4

用AI自动学奖励函数,提升6G卫星网络频谱效率

Enhancing Spectrum Efficiency in 6G Satellite Networks: A GAIL-Powered Policy Learning via Asynchronous Federated Inverse Reinforcement Learning

论文配图:Enhancing Spectrum Efficiency in 6G Satellite Networks: A GAIL-Powered Policy Learning via Asynchronous Federated Inverse Reinforcement Learning
图 1 · 摘自论文原文
  • 用生成对抗模仿学习自动学奖励,不用手动调参
  • 比传统方法快14.6%收敛,频谱效率显著提升
  • 适合研究6G卫星网络优化的学者和工程师

本文提出一种基于生成对抗模仿学习(GAIL)的策略学习方法,用于优化新型直连卫星网络(NTNs)中的波束成形、频谱分配和远端用户设备(RUE)关联。传统强化学习依赖人工设计奖励函数,需大量参数调优。为此,本文采用逆强化学习(IRL),利用GAIL框架自动学习奖励函数。通过异步联邦学习,多颗卫星可分布式协作生成最优策略。为应对该问题的非凸与NP难特性,结合多对一匹配理论与多智能体异步联邦逆强化学习(MA-AFIRL)框架,使智能体异步交互环境,提升训练效率与可扩展性。专家策略由鲸鱼优化算法(WOA)生成,为GAIL提供训练数据。仿真结果表明,所提方法在收敛速度与奖励值上相较传统强化学习提升14.6%,为6G NTNs优化建立新基准。

原文摘要 · Abstract (English)

In this paper, a novel generative adversarial imitation learning (GAIL)-powered policy learning approach is proposed for optimizing beamforming, spectrum allocation, and remote user equipment (RUE) association in NTNs. Traditional reinforcement learning (RL) methods for wireless network optimization often rely on manually designed reward functions, which can require extensive parameter tuning. To overcome these limitations, we employ inverse RL (IRL), specifically leveraging the GAIL framework, to automatically learn reward functions without manual design. We augment this framework with an asynchronous federated learning approach, enabling decentralized multi-satellite systems to collaboratively derive optimal policies. The proposed method aims to maximize spectrum efficiency (SE) while meeting minimum information rate requirements for RUEs. To address the non-convex, NP-hard nature of this problem, we combine the many-to-one matching theory with a multi-agent asynchronous federated IRL (MA-AFIRL) framework. This allows agents to learn through asynchronous environmental interactions, improving training efficiency and scalability. The expert policy is generated using the Whale optimization algorithm (WOA), providing data to train the automatic reward function within GAIL. Simulation results show that the proposed MA-AFIRL method outperforms traditional RL approaches, achieving a $14.6\%$ improvement in convergence and reward value. The novel GAIL-driven policy learning establishes a novel benchmark for 6G NTN optimization.

6G卫星强化学习频谱效率联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。