用语义引导强化学习,让无人机更智能地缓解城市车网断连问题。
Bridging Network Fragmentation: A Semantic-Augmented DRL Framework for UAV-aided VANETs
- 引入道路拓扑语义生成动作先验,指导无人机部署探索方向。
- 仅用28.6%训练轮次达成收敛,连接车辆数和连通区域大小提升超7%。
- 适合研究智能交通、无人机协同与强化学习应用的科研人员。
城市车联网因建筑遮挡与车辆移动导致网络频繁断连。无人机可作移动中继,但基于深度强化学习的部署常因缺乏道路拓扑引导而探索效率低。本文提出语义增强强化学习(SA-DRL),在道路拓扑上建模网络断连,并利用预训练大语言模型(LLM)从动态交通状态生成依赖拓扑的动作先验。所提语义增强PPO(SA-PPO)通过对数融合将该先验与PPO策略结合,引导探索聚焦高潜力交叉口,同时保留环境反馈的适应能力。基于真实城市轨迹的仿真表明,SA-PPO仅用28.6%的训练轮次即达到原始PPO的最终奖励水平;平均连通车辆数与连通组件大小分别提升7.9%和8.7%,无人机能耗降低21.3%。
原文摘要 · Abstract (English)
Urban Vehicular Ad-Hoc Networks (VANETs) can become fragmented because buildings obstruct wireless links and vehicle mobility continuously changes the network topology. Unmanned Aerial Vehicles (UAVs) can serve as mobile relays, but Deep Reinforcement Learning (DRL)-based deployment often suffers from inefficient exploration because it lacks road-topology guidance. To address this problem, we propose Semantic-Augmented DRL (SA-DRL), which models network fragmentation over the road topology and aligns a pretrained Large Language Model (LLM) to generate a topology-dependent action prior from dynamic traffic states. The resulting Semantic-Augmented PPO (SA-PPO) algorithm combines this prior with the PPO policy through Logit Fusion, guiding exploration toward promising intersections while retaining adaptation through environmental returns. Simulations driven by real-world urban trajectories show that SA-PPO reaches the final converged reward of Vanilla PPO using only 28.6% of its training episodes. It improves the average number of vehicles in connected components and the average connected-component size by 7.9% and 8.7%, respectively, while reducing UAV energy consumption by 21.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。