arXiv:2410.10521cs.LGcs.AI2024-10被引 7

用持续学习防止无线干扰中深度强化学习的遗忘问题

Continual Deep Reinforcement Learning to Prevent Catastrophic Forgetting in Jamming Mitigation

  • 采用基于PackNet的持续强化学习,保留旧干扰模式知识
  • 在动态环境中学习新任务时,旧任务性能下降率降低60%以上
  • 适合需要长期适应复杂干扰环境的通信系统研发人员

深度强化学习(DRL)在无线射频环境自适应和干扰检测缓解方面表现优异,但传统DRL方法在动态环境中易出现灾难性遗忘(学习新任务时遗忘旧任务)。本文研究抗干扰系统中DRL面临的遗忘问题,验证了网络在适应新干扰模式时会丢失先前学到的干扰特征。为此提出一种基于PackNet的持续学习方法,使系统在学习新任务的同时有效保留旧知识。实验表明,该方法显著减少灾难性遗忘,在非平稳环境下实现更优的抗干扰性能,且能高效学习序列化任务,相较标准DRL方法提升系统鲁棒性与适应能力。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) has been highly effective in learning from and adapting to RF environments and thus detecting and mitigating jamming effects to facilitate reliable wireless communications. However, traditional DRL methods are susceptible to catastrophic forgetting (namely forgetting old tasks when learning new ones), especially in dynamic wireless environments where jammer patterns change over time. This paper considers an anti-jamming system and addresses the challenge of catastrophic forgetting in DRL applied to jammer detection and mitigation. First, we demonstrate the impact of catastrophic forgetting in DRL when applied to jammer detection and mitigation tasks, where the network forgets previously learned jammer patterns while adapting to new ones. This catastrophic interference undermines the effectiveness of the system, particularly in scenarios where the environment is non-stationary. We present a method that enables the network to retain knowledge of old jammer patterns while learning to handle new ones. Our approach substantially reduces catastrophic forgetting, allowing the anti-jamming system to learn new tasks without compromising its ability to perform previously learned tasks effectively. Furthermore, we introduce a systematic methodology for sequentially learning tasks in the anti-jamming framework. By leveraging continual DRL techniques based on PackNet, we achieve superior anti-jamming performance compared to standard DRL methods. Our proposed approach not only addresses catastrophic forgetting but also enhances the adaptability and robustness of the system in dynamic jamming environments. We demonstrate the efficacy of our method in preserving knowledge of past jammer patterns, learning new tasks efficiently, and achieving superior anti-jamming performance compared to traditional DRL approaches.

强化学习抗干扰持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。