arXiv:2511.17931cs.ITcs.LG2025-11中稿 · IEEE Trans

用强化学习优化上行载波聚合,自动避开自干扰,提升用户速率。

A Reinforcement Learning Framework for Resource Allocation in Uplink Carrier Aggregation in the Presence of Self Interference

  • 设计复合动作的强化学习算法,同时决策开哪些载波和分配多少功率。
  • 相比传统方法,上行总吞吐量提升明显,且能适应自干扰存在与否的场景。
  • 适合研究无线资源调度或通信系统优化的工程师与研究人员。

载波聚合(CA)通过合并多个载波提升用户数据速率。在上行链路中,受限于功率的用户需高效分配可用功率到指定载波。若上行载波的谐波落在自身下行频段,将引发自干扰导致下行接收灵敏度下降。本文将上行载波聚合建模为带非线性自干扰约束的最优资源分配问题,涉及离散变量(激活哪些载波)和连续变量(功率分配),且环境动态变化,传统方法难以求解。为此,提出基于复合动作演员-评论家(CA2C)的强化学习框架,并设计关键奖励函数以有效处理自干扰。该方案可在线学习并动态选择合适载波与功率配置。数值结果表明,所提方法相比基准方案显著提升总吞吐量;且奖励函数使算法能在有无自干扰环境下均有效适应。

原文摘要 · Abstract (English)

Carrier aggregation (CA) is a technique that allows mobile networks to combine multiple carriers to increase user data rate. On the uplink, for power constrained users, this translates to the need for an efficient resource allocation scheme, where each user distributes its available power among its assigned uplink carriers. Choosing a good set of carriers and allocating appropriate power on the carriers is important. If the carrier allocation on the uplink is such that a harmonic of a user's uplink carrier falls on the downlink frequency of that user, it leads to a self coupling-induced sensitivity degradation of that user's downlink receiver. In this paper, we model the uplink carrier aggregation problem as an optimal resource allocation problem with the associated constraints of non-linearities induced self interference (SI). This involves optimization over a discrete variable (which carriers need to be turned on) and a continuous variable (what power needs to be allocated on the selected carriers) in dynamic environments, a problem which is hard to solve using traditional methods owing to the mixed nature of the optimization variables and the additional need to consider the SI constraint. We adopt a reinforcement learning (RL) framework involving a compound-action actor-critic (CA2C) algorithm for the uplink carrier aggregation problem. We propose a novel reward function that is critical for enabling the proposed CA2C algorithm to efficiently handle SI. The CA2C algorithm along with the proposed reward function learns to assign and activate suitable carriers in an online fashion. Numerical results demonstrate that the proposed RL based scheme is able to achieve higher sum throughputs compared to naive schemes. The results also demonstrate that the proposed reward function allows the CA2C algorithm to adapt the optimization both in the presence and absence of SI.

强化学习载波聚合自干扰资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。