改进的噪声DQN让无人机在复杂环境中更快更稳地规划飞行路径。
UAV Trajectory Optimization via Improved Noisy Deep Q-Network
- 结合残差噪声层与自适应噪声调度,增强探索能力。
- 训练收敛更快,奖励提升最高达40%,任务步数减少28%。
- 适合需要高效稳定强化学习的无人机路径规划场景。
本文提出一种改进的噪声深度Q网络(Noisy DQN),以提升无人飞行器(UAV)在模拟环境中应用深度强化学习时的探索能力和训练稳定性。该方法通过将残差噪声线性层与自适应噪声调度机制结合,增强探索能力;同时利用平滑损失和软目标网络更新提升训练稳定性。实验表明,所提模型在15×15网格导航环境中的任务28上,收敛速度更快,奖励最高提升40%,并迅速达到任务所需的最少步数。结果表明,对NoisyNet网络结构、探索控制及训练稳定性的综合改进,显著提升了深度Q学习的效率与可靠性。
原文摘要 · Abstract (English)
This paper proposes an Improved Noisy Deep Q-Network (Noisy DQN) to enhance the exploration and stability of Unmanned Aerial Vehicle (UAV) when applying deep reinforcement learning in simulated environments. This method enhances the exploration ability by combining the residual NoisyLinear layer with an adaptive noise scheduling mechanism, while improving training stability through smooth loss and soft target network updates. Experiments show that the proposed model achieves faster convergence and up to $+40$ higher rewards compared to standard DQN and quickly reach to the minimum number of steps required for the task 28 in the 15 * 15 grid navigation environment set up. The results show that our comprehensive improvements to the network structure of NoisyNet, exploration control, and training stability contribute to enhancing the efficiency and reliability of deep Q-learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。