arXiv:2605.00457cs.NIcs.LG2026-05

用强化学习动态调整NR-U传输时长,平衡网络公平与效率。

Utility-Aware DRL-Based TXOP Adaptation for NR-U and Wi-Fi Coexistence Networks

  • 基于深度强化学习,动态优化NR-U的传输机会时长。
  • 在严格公平下公平性指数超0.9,吞吐量提升达177.6%。
  • 支持多种策略配置,适合不同场景下的共存需求。

NR-U与Wi-Fi在非授权频谱中的共存带来了资源管理挑战,异构接入机制导致频谱利用不均和Wi-Fi性能严重下降。本文提出一种面向效用的深度强化学习(DRL)框架,用于NR-U/Wi-Fi共存网络中的自适应传输机会(TXOP)控制。将共存过程建模为马尔可夫决策过程(MDP),将NR-U的TXOP时长作为可调控变量以调节信道占用。采用深度Q网络(DQN)通过与环境在线交互学习自适应的TXOP控制策略。框架核心在于可配置的奖励与准则设计,实现对公平性-效率-效用权衡的显式控制。设计了三种运行策略:绝对公平、适度公平与效用导向的适度公平,以表征不同共存状态。仿真结果表明,在严格公平控制下,该框架的Jain公平性指数超过0.9;相比绝对公平策略,适度公平策略使总吞吐量提升68.22%;而效用导向策略在所采用的效用评估指标下实现177.6%的提升。结果表明,该效用感知的DRL框架为异构非授权共存网络提供了有效且灵活的自适应TXOP控制与权衡管理方案。

原文摘要 · Abstract (English)

The coexistence of NR-U and Wi-Fi in the unlicensed spectrum introduces a challenging resource management problem, where heterogeneous channel access mechanisms can lead to unbalanced spectrum utilization and severe Wi-Fi performance degradation. To address this issue, this paper proposes a utility-aware deep reinforcement learning (DRL) framework for adaptive transmission opportunity (TXOP) control in NR-U/Wi-Fi coexistence networks. The coexistence process is formulated as a Markov decision process (MDP), in which the NR-U TXOP duration is treated as a controllable variable for regulating post-access channel occupancy. A deep Q-network (DQN) is then employed to learn adaptive TXOP control policies through online interaction with the coexistence environment. A key feature of the proposed framework is the integration of a configurable reward and criterion design, which enables explicit control of the fairness-efficiency-utility tradeoff. Three operating policies are developed, namely absolute fairness, moderate fairness, and utility-oriented moderate fairness, to characterize different coexistence operating points. Simulation results show that the proposed framework achieves a Jain fairness index above 0.9 under strict fairness control. Compared with the absolute fairness policy, the moderate fairness policy improves aggregate throughput by 68.22%, while the utility-oriented policy achieves a 177.6% improvement under the adopted utility evaluation metric. These results demonstrate that the proposed utility-aware DRL framework provides an effective and flexible solution for adaptive TXOP control and tradeoff management in heterogeneous unlicensed coexistence networks.

强化学习频谱共存通信系统动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。