arXiv:2410.21521cs.LGcs.AI2024-10中稿 · IEEE CCNC 2025被引 3

构建多智能体强化学习环境,模拟复杂无线频谱场景。

A Multi-Agent Reinforcement Learning Testbed for Cognitive Radio Applications

  • 基于Ray RLlib扩展单智能体环境,支持多智能体协同与竞争训练。
  • 可模拟多种真实频谱场景,验证算法在复杂环境下的性能表现。
  • 适合研究无线资源分配、频谱共享等场景的科研人员使用。

技术趋势表明,射频强化学习(RFRL)将在未来无线通信系统中发挥重要作用,应用涵盖军事通信干扰到提升WiFi网络性能。在部署算法前,需在仿真环境中进行训练以确保性能。为此,我们此前开发了RFRL Gym——一个基于OpenAI Gym框架的标准化工具,支持无线通信领域强化学习算法的开发与测试,具备可定制的射频频谱仿真场景。然而,原版RFRL Gym仅支持单智能体训练,难以反映现实世界中多智能体共存、合作或竞争的复杂情况。为解决此问题,本工作通过集成Ray RLlib,新增多智能体强化学习(MARL)功能,实现了对多智能体算法的训练与评估能力。本文介绍更新后的RFRL Gym架构,对比现有资源,突出其关键改进与重构,并展示在多种射频场景中测试MARL算法的结果,同时讨论未来扩展方向。

原文摘要 · Abstract (English)

Technological trends show that Radio Frequency Reinforcement Learning (RFRL) will play a prominent role in the wireless communication systems of the future. Applications of RFRL range from military communications jamming to enhancing WiFi networks. Before deploying algorithms for these purposes, they must be trained in a simulation environment to ensure adequate performance. For this reason, we previously created the RFRL Gym: a standardized, accessible tool for the development and testing of reinforcement learning (RL) algorithms in the wireless communications space. This environment leveraged the OpenAI Gym framework and featured customizable simulation scenarios within the RF spectrum. However, the RFRL Gym was limited to training a single RL agent per simulation; this is not ideal, as most real-world RF scenarios will contain multiple intelligent agents in cooperative, competitive, or mixed settings, which is a natural consequence of spectrum congestion. Therefore, through integration with Ray RLlib, multi-agent reinforcement learning (MARL) functionality for training and assessment has been added to the RFRL Gym, making it even more of a robust tool for RF spectrum simulation. This paper provides an overview of the updated RFRL Gym environment. In this work, the general framework of the tool is described relative to comparable existing resources, highlighting the significant additions and refactoring we have applied to the Gym. Afterward, results from testing various RF scenarios in the MARL environment and future additions are discussed.

强化学习多智能体频谱管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。