用强化学习优化雷达资源分配,兼顾新目标探测与旧目标跟踪。
Multi-Objective Reinforcement Learning for Cognitive Radar Resource Management
- 将雷达时隙分配建模为多目标优化问题,用深度强化学习求解帕累托最优。
- SAC算法比DDPG更稳定、样本效率更高,在多种场景下表现优异。
- 结合NSGA-II算法估算最优解边界,为系统设计提供参考,适合雷达算法研究者。
多任务认知雷达系统中的时隙分配问题,核心在于新出现目标的扫描与已有目标跟踪之间的权衡。本文将其建模为多目标优化问题,采用深度强化学习方法寻找帕累托最优解,并对比了深度确定性策略梯度(DDPG)与软演员-评论家(SAC)算法。实验结果表明,两种算法均能有效适应不同场景,其中SAC在稳定性与样本效率方面优于DDPG。此外,本文还采用NSGA-II算法估计该问题帕累托前沿的上界。本研究推动了动态环境中可高效自适应的多目标认知雷达系统的发展。
原文摘要 · Abstract (English)
The time allocation problem in multi-function cognitive radar systems focuses on the trade-off between scanning for newly emerging targets and tracking the previously detected targets. We formulate this as a multi-objective optimization problem and employ deep reinforcement learning to find Pareto-optimal solutions and compare deep deterministic policy gradient (DDPG) and soft actor-critic (SAC) algorithms. Our results demonstrate the effectiveness of both algorithms in adapting to various scenarios, with SAC showing improved stability and sample efficiency compared to DDPG. We further employ the NSGA-II algorithm to estimate an upper bound on the Pareto front of the considered problem. This work contributes to the development of more efficient and adaptive cognitive radar systems capable of balancing multiple competing objectives in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。