用因果感知强化学习优化6G通信感知一体化系统波束成形
Causality-Driven Reinforcement Learning for Joint Communication and Sensing
- 通过状态相关动作维度选择实现因果关系发现
- 相比基线方法提升波束成形增益,降低训练开销
- 适合6G车联网等对实时性要求高的场景
下一代无线网络(6G及以后)旨在融合通信与感知,以克服干扰、提升频谱效率并降低硬件与功耗。基于大规模多输入多输出(mMIMO)的通信与感知一体化(JCAS)系统为自动驾驶等应用提供了精准环境感知与低时延通信支持。现有研究使用强化学习(RL)进行mMIMO波束成形,但波束成形的动作空间庞大,导致学习效率低下,且未考虑动作与奖励之间的因果关系,所有动作被同等对待。本文提出一种因果感知的强化学习代理,在训练阶段可干预并发现mMIMO-JCAS环境中的因果关系。采用状态依赖的动作维度选择策略实现因果发现。在多种JCAS场景下的评估表明,所提框架在波束成形增益方面优于基线方法。
原文摘要 · Abstract (English)
The next-generation wireless network, 6G and beyond, envisions to integrate communication and sensing to overcome interference, improve spectrum efficiency, and reduce hardware and power consumption. Massive Multiple-Input Multiple Output (mMIMO)-based Joint Communication and Sensing (JCAS) systems realize this integration for 6G applications such as autonomous driving, as it requires accurate environmental sensing and time-critical communication with neighboring vehicles. Reinforcement Learning (RL) is used for mMIMO antenna beamforming in the existing literature. However, the huge search space for actions associated with antenna beamforming causes the learning process for the RL agent to be inefficient due to high beam training overhead. The learning process does not consider the causal relationship between action space and the reward, and gives all actions equal importance. In this work, we explore a causally-aware RL agent which can intervene and discover causal relationships for mMIMO-based JCAS environments, during the training phase. We use a state dependent action dimension selection strategy to realize causal discovery for RL-based JCAS. Evaluation of the causally-aware RL framework in different JCAS scenarios shows the benefit of our proposed framework over baseline methods in terms of the beamforming gain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。