自动驾驶中用强化学习动态调整毫米波通信,提升感知与传输稳定性。
Joint Adaptive OFDM and Reinforcement Learning Design for Autonomous Vehicles: Leveraging Age of Updates
- 结合队列与信道状态,用强化学习动态调节OFDM参数。
- 采用更新时效性设计奖励函数,降低丢包率并提升速度分辨率。
- 适用于高动态自动驾驶场景,兼顾通信与感知需求。
基于毫米波(mmWave)的正交频分复用(OFDM)是实现高分辨率感知与高速数据传输的优选方案。传统方法采用静态配置,预设帧内符号数和通信时隙内帧数,但环境动态且无线信道易受干扰,尤其在自动驾驶(AV)系统中尤为突出。本文提出一种融合感知与通信(ISAC)的自适应框架:自动驾驶车辆利用队列状态信息(QSI)与信道状态信息(CSI),结合强化学习技术,实现通信与感知的协同优化。目标包括维持与其他车辆的稳定通信链路,以及高分辨率估计周围物体的速度。通信性能通过队列状态、有效数据速率和丢包率评估;感知效果以速度分辨率为指标。系统采用自适应OFDM实现动态调制,并设计基于更新时效性的奖励函数,以缓解通信缓冲压力并提升感知精度。使用优势演员-评论家(A2C)和近端策略优化(PPO)算法进行验证,仿真结果表明,相比现有设计,本方案在通信稳定性与感知精度上均表现更优。
原文摘要 · Abstract (English)
Millimeter wave (mmWave)-based orthogonal frequency-division multiplexing (OFDM) stands out as a suitable alternative for high-resolution sensing and high-speed data transmission. To meet communication and sensing requirements, many works propose a static configuration where the wave's hyperparameters such as the number of symbols in a frame and the number of frames in a communication slot are already predefined. However, two facts oblige us to redefine the problem, (1) the environment is often dynamic and uncertain, and (2) mmWave is severely impacted by wireless environments. A striking example where this challenge is very prominent is autonomous vehicle (AV). Such a system leverages integrated sensing and communication (ISAC) using mmWave to manage data transmission and the dynamism of the environment. In this work, we consider an autonomous vehicle network where an AV utilizes its queue state information (QSI) and channel state information (CSI) in conjunction with reinforcement learning techniques to manage communication and sensing. This enables the AV to achieve two primary objectives: establishing a stable communication link with other AVs and accurately estimating the velocities of surrounding objects with high resolution. The communication performance is therefore evaluated based on the queue state, the effective data rate, and the discarded packets rate. In contrast, the effectiveness of the sensing is assessed using the velocity resolution. In addition, we exploit adaptive OFDM techniques for dynamic modulation, and we suggest a reward function that leverages the age of updates to handle the communication buffer and improve sensing. The system is validated using advantage actor-critic (A2C) and proximal policy optimization (PPO). Furthermore, we compare our solution with the existing design and demonstrate its superior performance by computer simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。