arXiv:2605.00059cs.ROcs.AI2026-05

动态预测障碍物轨迹,让无人机更安全地避障飞行

Dynamic-TD3: A Novel Algorithm for UAV Path Planning with Dynamic Obstacle Trajectory Prediction

论文配图:Dynamic-TD3: A Novel Algorithm for UAV Path Planning with Dynamic Obstacle Trajectory Prediction
图 1 · 摘自论文原文
  • 用动态轨迹预测机制捕捉障碍物长期运动意图
  • 在强干扰下仍能实现零碰撞、能耗降低18%
  • 适合高动态复杂环境中的无人机自主导航

深度强化学习在复杂高风险环境中的无人机自主导航中应用广泛。但实际部署面临安全与探索的矛盾:软惩罚机制易导致冒险试错,多数约束方法在传感器噪声和意图不确定性下性能下降。本文提出Dynamic-TD3,一种基于物理增强的框架,将导航建模为带约束的马尔可夫决策过程(CMDP),通过自适应轨迹关系演化机制(ATREM)捕捉长程意图,并采用物理感知门控卡尔曼滤波器(PAG-KF)抑制非平稳观测噪声。该状态表示驱动双目标策略,利用拉格朗日松弛平衡任务效率与硬性安全约束。在激进动态威胁实验中,该方法展现出优异的避碰性能,能量消耗减少18%,飞行轨迹更平滑。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL) finds extensive application in autonomous drone navigation within complex, high-risk environments. However, its practical deployment faces a safety-exploration dilemma: soft penalty mechanisms encourage risky trial-and-error, while most constraint-based methods suffer degraded performance under sensor noise and intent uncertainty. We propose Dynamic-TD3, a physically enhanced framework that enforces strict safety constraints while maintaining maneuverability by modeling navigation as a Constrained Markov Decision Process (CMDP). This framework integrates an Adaptive Trajectory Relational Evolution Mechanism (ATREM) to capture long-range intentions and employs a Physically Aware Gated Kalman Filter (PAG-KF) to mitigate non-stationary observation noise. The resulting state representation drives a dual-criterion policy that balances mission efficiency against hard safety constraints via Lagrangian relaxation. In experiments with aggressive dynamic threats, this approach demonstrates superior collision avoidance performance, reduced energy consumption, and smoother flight trajectories.

无人机导航强化学习避障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。