让无人机在质量变化和电机故障下仍能高速精准飞行。
MAVEN: A Meta-Reinforcement Learning Framework for Varying-Dynamics Expertise in Agile Quadrotor Maneuvers
- 用历史数据预测飞行器动态,实现自适应控制
- 支持66.7%质量变化和70%单电机失效下的稳定飞行
- 一小时完成训练,真实无人机可直接部署
强化学习(RL)已成功用于实现四旋翼无人机的在线敏捷导航。然而,传统RL训练的策略在面对显著动态变化时通常无法泛化,缺乏适应能力。本文提出MAVEN,一种元强化学习框架,使单一策略可在多种四旋翼动力学条件下实现鲁棒的端到端导航。该方法引入新颖的预测性上下文编码器,从交互历史中学习系统动态的潜在表示。我们在两种挑战性场景下验证了该方法:四旋翼质量大幅变化以及严重单旋翼推力损失。利用GPU向量化模拟器,将任务分布于数千个并行环境,克服了元强化学习长期训练的问题,实现一小时内收敛。通过仿真与真实世界中的大量实验,验证了MAVEN在适应性和敏捷性上的优越表现。策略成功实现零样本仿真到现实的迁移,在质量变化高达66.7%、单旋翼推力损失达70%的情况下,仍能完成高速机动任务。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has emerged as a powerful paradigm for achieving online agile navigation with quadrotors. Despite this success, policies trained via standard RL typically fail to generalize across significant dynamic variations, exhibiting a critical lack of adaptability. This work introduces MAVEN, a meta-RL framework that enables a single policy to achieve robust end-to-end navigation across a wide range of quadrotor dynamics. Our approach features a novel predictive context encoder, which learns to infer a latent representation of the system dynamics from interaction history. We demonstrate our method in agile waypoint traversal tasks under two challenging scenarios: large variations in quadrotor mass and severe single-rotor thrust loss. We leverage a GPU-vectorized simulator to distribute tasks across thousands of parallel environments, overcoming the long training times of meta-RL to converge in less than an hour. Through extensive experiments in both simulation and the real world, we validate that MAVEN achieves superior adaptation and agility. The policy successfully executes zero-shot sim-to-real transfer, demonstrating robust online adaptation by performing high-speed maneuvers despite mass variations of up to 66.7% and single-rotor thrust losses as severe as 70%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。