arXiv:2604.21130cs.RO2026-04

提出自预测模型AmelPredSto,提升无人机目标导航的样本效率。

Self-Predictive Representation for Autonomous UAV Object-Goal Navigation

论文配图:Self-Predictive Representation for Autonomous UAV Object-Goal Navigation
图 1 · 摘自论文原文
  • 设计自预测感知模块,让无人机自主学习环境表征。
  • 在3D目标导航任务中,结合强化学习使样本效率显著提升。
  • 适合关注无人机自主导航与表征学习的研究者。

自主无人机(UAV)在航拍监视、搜救、农业和配送等领域广泛应用,其自主能力可在大范围开放空间中运行。强化学习(RL)使无人机能够学习复杂导航策略,但数据样本利用效率低仍是主要挑战。在物体目标导航(OGN)场景中,目标识别成为额外难点。现有方法多依赖相对或绝对坐标从起点到达预设位置,而非直接寻找目标。本研究针对3D OGN问题中的样本效率问题,将未知目标位置设定形式化为马尔可夫决策过程。通过实验分析不同状态表征学习(SRL)方法与无模型强化学习算法在自主导航系统中的协同作用。主要贡献是开发了新型自预测模型AmelPred。实证结果表明,其随机版本AmelPredSto在与演员-评论家强化学习算法结合时表现最佳,显著提升了在解决OGN问题中的算法效率。

原文摘要 · Abstract (English)

Autonomous Unmanned Aerial Vehicles (UAVs) have revolutionized industries through their versatility with applications including aerial surveillance, search and rescue, agriculture, and delivery. Their autonomous capabilities offer unique advantages, such as operating in large open space environments. Reinforcement Learning (RL) empowers UAVs to learn intricate navigation policies, enabling them to optimize flight behavior autonomously. However, one of its main challenge is the inefficiency in using data sample to achieve a good policy. In object-goal navigation (OGN) settings, target recognition arises as an extra challenge. Most UAV-related approaches use relative or absolute coordinates to move from an initial position to a predefined location, rather than to find the target directly. This study addresses the data sample efficiency issue in solving a 3D OGN problem, in addition to, the formalization of the unknown target location setting as a Markov decision process. Experiments are conducted to analyze the interplay of different state representation learning (SRL) methods for perception with a model-free RL algorithm for planning in an autonomous navigation system. The main contribution of this study is the development of the perception module, featuring a novel self-predictive model named AmelPred. Empirical results demonstrate that its stochastic version, AmelPredSto, is the best-performing SRL model when combined with actor-critic RL algorithms. The obtained results show substantial improvement in RL algorithms' efficiency by using AmelPredSto in solving the OGN problem.

无人机导航强化学习表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。