用量子强化学习提升蓝牙设备共享频谱时的通信效率。
Dynamic Spectrum Access for Ambient Backscatter Communication-assisted D2D Systems with Quantum Reinforcement Learning
- 用量子电路替代深度网络,实现更快策略学习。
- 频谱繁忙时平均吞吐量提升显著,收敛速度提高3倍以上。
- 适合资源受限的物联网场景,对低延迟通信研究者有参考价值。
频谱共享是设备到设备(D2D)通信中的核心挑战。随着移动设备激增,无线频谱日益紧张,导致D2D通信频谱效率低下。为此,本文将环境背散射通信技术引入D2D设备,使其在共享频谱被占用时,可通过背散射环境射频信号传输数据。为获取最优频谱接入策略(空闲、主动发送或背散射),以最大化D2D用户平均吞吐量,可采用深度强化学习(DRL)。但传统DRL因维度灾难和复杂网络结构导致训练时间长。为此,本文提出一种新型量子强化学习(QRL)算法,利用量子叠加与纠缠特性,在更少参数下实现更快收敛。具体地,采用可调量子电路替代传统深度神经网络来逼近最优策略。大量仿真表明,该方案不仅显著提升共享频谱繁忙时的D2D平均吞吐量,且在收敛速度与学习复杂度方面均优于现有DRL方法。
原文摘要 · Abstract (English)
Spectrum access is an essential problem in device-to-device (D2D) communications. However, with the recent growth in the number of mobile devices, the wireless spectrum is becoming scarce, resulting in low spectral efficiency for D2D communications. To address this problem, this paper aims to integrate the ambient backscatter communication technology into D2D devices to allow them to backscatter ambient RF signals to transmit their data when the shared spectrum is occupied by mobile users. To obtain the optimal spectrum access policy, i.e., stay idle or access the shared spectrum and perform active transmissions or backscattering ambient RF signals for transmissions, to maximize the average throughput for D2D users, deep reinforcement learning (DRL) can be adopted. However, DRL-based solutions may require long training time due to the curse of dimensionality issue as well as complex deep neural network architectures. For that, we develop a novel quantum reinforcement learning (RL) algorithm that can achieve a faster convergence rate with fewer training parameters compared to DRL thanks to the quantum superposition and quantum entanglement principles. Specifically, instead of using conventional deep neural networks, the proposed quantum RL algorithm uses a parametrized quantum circuit to approximate an optimal policy. Extensive simulations then demonstrate that the proposed solution not only can significantly improve the average throughput of D2D devices when the shared spectrum is busy but also can achieve much better performance in terms of convergence rate and learning complexity compared to existing DRL-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。