arXiv:2412.20098cs.ROcs.AI2024-12

针对大中型无人机动态避障与地形跟随,提出融合流场与强化学习的实时路径规划方法。

RFPPO: Motion Dynamic RRT based Fluid Field - PPO for Dynamic TF/TA Routing Planning

  • 基于扰动流场与人工势场重构状态动作空间,建模飞机动力学特性。
  • 在真实数字高程数据上实现无全局规划的长距离无碰撞飞行。
  • 适合需要实时性与动态约束满足的大中型固定翼无人机任务。

现有局部动态路径规划算法在应用于大中型固定翼无人机的地形跟随/避障或动态障碍物规避时,难以同时满足实时性、长距离规划及飞行器动态约束要求。为此,本文提出一种基于运动动态RRT的流场-近端策略优化(Motion Dynamic RRT based Fluid Field - PPO)算法,用于动态地形跟随/避障路径规划。首先,利用扰动流场与人工势场算法重新设计近端策略梯度(PPO)算法的动作与状态空间,建立飞机动力学模型,并据此设计状态转移过程;此外,设计奖励函数以鼓励避障、地形跟随、地形规避与安全飞行等策略。在真实数字高程模型(DEM)数据上的实验结果表明,该算法可在无需预先全局规划的情况下,完成符合动态约束的无碰撞长距离飞行任务。

原文摘要 · Abstract (English)

Existing local dynamic route planning algorithms, when directly applied to terrain following/terrain avoidance, or dynamic obstacle avoidance for large and medium-sized fixed-wing aircraft, fail to simultaneously meet the requirements of real-time performance, long-distance planning, and the dynamic constraints of large and medium-sized aircraft. To deal with this issue, this paper proposes the Motion Dynamic RRT based Fluid Field - PPO for dynamic TF/TA routing planning. Firstly, the action and state spaces of the proximal policy gradient algorithm are redesigned using disturbance flow fields and artificial potential field algorithms, establishing an aircraft dynamics model, and designing a state transition process based on this model. Additionally, a reward function is designed to encourage strategies for obstacle avoidance, terrain following, terrain avoidance, and safe flight. Experimental results on real DEM data demonstrate that our algorithm can complete long-distance flight tasks through collision-free trajectory planning that complies with dynamic constraints, without the need for prior global planning.

路径规划强化学习无人机地形跟随

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。