arXiv:2502.01268cs.ROcs.AI2025-02被引 3

用少量离线数据让无人机快速适应新环境,同时降低信息延迟和能耗。

Resilient UAV Trajectory Planning via Few-Shot Meta-Offline Reinforcement Learning

  • 结合保守Q学习与元学习,实现无需在线交互的离线强化学习。
  • 仅用少量数据点即优化无人机轨迹,收敛速度优于传统方法。
  • 能应对突发网络故障,适合动态变化的无线通信场景。

强化学习在未来的5G演进和6G系统中具有巨大潜力,其优势在于复杂高维无线环境中无需建模的鲁棒决策能力。然而,现有多数RL框架依赖于与环境的在线交互,这在安全和成本方面可能不可行。此外,在动态或新环境下的可扩展性不足也是问题。本文提出一种新型、鲁棒的少样本元离线强化学习算法,结合使用保守Q学习(CQL)的离线强化学习与基于模型无关元学习(MAML)的元学习。该算法可在不进行任何在线交互的情况下,利用静态离线数据集训练强化学习模型,并借助MAML实现对新未见环境的可扩展性。我们以优化无人机(UAV)的轨迹与调度策略为目标,最小化有限功率设备的年龄信息(AoI)和传输功率。数值结果表明,所提算法收敛速度优于深度Q网络和CQL等基线方案;且是唯一能在极少数数据点下实现最优联合AoI与传输功率的算法,对前所未有的环境变化具有抗扰能力。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has been a promising essence in future 5G-beyond and 6G systems. Its main advantage lies in its robust model-free decision-making in complex and large-dimension wireless environments. However, most existing RL frameworks rely on online interaction with the environment, which might not be feasible due to safety and cost concerns. Another problem with online RL is the lack of scalability of the designed algorithm with dynamic or new environments. This work proposes a novel, resilient, few-shot meta-offline RL algorithm combining offline RL using conservative Q-learning (CQL) and meta-learning using model-agnostic meta-learning (MAML). The proposed algorithm can train RL models using static offline datasets without any online interaction with the environments. In addition, with the aid of MAML, the proposed model can be scaled up to new unseen environments. We showcase the proposed algorithm for optimizing an unmanned aerial vehicle (UAV) 's trajectory and scheduling policy to minimize the age-of-information (AoI) and transmission power of limited-power devices. Numerical results show that the proposed few-shot meta-offline RL algorithm converges faster than baseline schemes, such as deep Q-networks and CQL. In addition, it is the only algorithm that can achieve optimal joint AoI and transmission power using an offline dataset with few shots of data points and is resilient to network failures due to unprecedented environmental changes.

强化学习无人机离线学习元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。