arXiv:2501.14377cs.RO2025-01中稿 · ICRA被引 17

用视觉像素直接控制无人机竞速,实现9米/秒高速飞行。

Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight

  • 基于模型强化学习,仅用摄像头像素训练飞行策略
  • 在仿真与真实飞行中均达9米/秒速度,无需人为设计视角奖励
  • 自动聚焦纹理丰富的门区,适合追求端到端视觉控制的研究者

自主无人机竞速已成为测试学习、感知、规划与控制极限的挑战性机器人基准。专家人类飞行员能通过单目相机的像素直接映射为控制指令完成竞速。现有基于直接像素-指令策略的自主竞速方法多依赖简化观测空间的中间表示,或通过模仿学习进行大量预训练。本文采用DreamerV3,在仅使用像素作为观测的情况下,训练出具备敏捷飞行能力的视觉-运动策略。相比PPO或SAC等无模型方法在该任务中样本效率低且表现不佳,本方法可从像素中直接习得竞速技能。值得注意的是,系统自发产生面向纹理丰富门区的感知行为,无需人工设计视角奖励项。实验表明,该方法在仿真和真实世界(硬件在环、渲染图像观测)中均可部署于真实四轴无人机,最高飞行速度达9米/秒。这些结果推进了基于像素的自主飞行技术,证明模型增强强化学习为现实机器人研究提供了可行路径。

原文摘要 · Abstract (English)

Autonomous drone racing has risen as a challenging robotic benchmark for testing the limits of learning, perception, planning, and control. Expert human pilots are able to fly a drone through a race track by mapping pixels from a single camera directly to control commands. Recent works in autonomous drone racing attempting direct pixel-to-commands control policies have relied on either intermediate representations that simplify the observation space or performed extensive bootstrapping using Imitation Learning (IL). This paper leverages DreamerV3 to train visuomotor policies capable of agile flight through a racetrack using only pixels as observations. In contrast to model-free methods like PPO or SAC, which are sample-inefficient and struggle in this setting, our approach acquires drone racing skills from pixels. Notably, a perception-aware behaviour of actively steering the camera toward texture-rich gate regions emerges without the need of handcrafted reward terms for the viewing direction. Our experiments show in both, simulation and real-world flight using a hardware-in-the-loop setup with rendered image observations, how the proposed approach can be deployed on real quadrotors at speeds of up to 9 m/s. These results advance the state of pixel-based autonomous flight and demonstrate that MBRL offers a promising path for real-world robotics research.

无人机竞速视觉控制模型强化学习端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。