让无人机快速学会在陌生城市导航并高效迁移
UAS Visual Navigation in Large and Unseen Environments via a Meta Agent
- 用分层课程训练+元强化学习,提升泛化能力
- 新算法ISAR使收敛速度比传统方法快得多
- 适合需要跨环境快速适应的无人机系统
本文旨在开发一种方法,使无人飞行系统(UAS)能够高效学习在大规模城市环境中导航,并将其所学知识迁移至全新环境。为此,我们提出一种元课程训练方案:首先通过元训练使智能体学习一个通用策略以跨任务泛化,随后在下游任务上进行微调。训练过程采用分层组织,引导智能体从粗略到精细逐步逼近目标任务。此外,引入增量自适应强化学习(ISAR)算法,结合增量学习与元强化学习思想。相较于传统强化学习专注于单一任务策略,元强化学习旨在学习具备快速迁移能力的策略,但训练耗时较长;而ISAR算法显著加速收敛。我们在模拟环境中评估该方法,结果表明,结合此训练范式与ISAR算法,能大幅提升城市级导航的收敛速度及在新环境中的适应能力。
原文摘要 · Abstract (English)
The aim of this work is to develop an approach that enables Unmanned Aerial System (UAS) to efficiently learn to navigate in large-scale urban environments and transfer their acquired expertise to novel environments. To achieve this, we propose a meta-curriculum training scheme. First, meta-training allows the agent to learn a master policy to generalize across tasks. The resulting model is then fine-tuned on the downstream tasks. We organize the training curriculum in a hierarchical manner such that the agent is guided from coarse to fine towards the target task. In addition, we introduce Incremental Self-Adaptive Reinforcement learning (ISAR), an algorithm that combines the ideas of incremental learning and meta-reinforcement learning (MRL). In contrast to traditional reinforcement learning (RL), which focuses on acquiring a policy for a specific task, MRL aims to learn a policy with fast transfer ability to novel tasks. However, the MRL training process is time consuming, whereas our proposed ISAR algorithm achieves faster convergence than the conventional MRL algorithm. We evaluate the proposed methodologies in simulated environments and demonstrate that using this training philosophy in conjunction with the ISAR algorithm significantly improves the convergence speed for navigation in large-scale cities and the adaptation proficiency in novel environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。