arXiv:2409.00895cs.RO2024-09ICRA被引 25

纯数据驱动方法让无人机通过狭窄缝隙,直接从像素和感知输出控制指令。

Whole-Body Control Through Narrow Gaps From Pixels To Action

  • 神经网络直接从像素与本体感知映射到低层控制命令。
  • 可在不同几何缝隙中实现急转弯姿态控制,成功穿越狭窄空间。
  • 适合研究无人机自主飞行、强化学习与视觉控制融合的场景。

在环境中穿越与身体尺寸相当的狭窄缝隙是欠驱动多旋翼无人机最具挑战性的飞行任务之一。本文探索了一种纯粹的数据驱动方法,在仿真中掌握此项飞行技能,即利用神经网络将像素信息与本体感知直接映射为连续的低层控制命令。该学习策略实现了在不同几何结构缝隙中的全身控制,可应对剧烈的姿态变化(如接近垂直的滚转角)。策略通过逐级无模型强化学习(RL)与在线观测空间蒸馏实现:强化学习阶段使用(虚拟)点云表示缝隙边缘以支持可扩展仿真,随后将其蒸馏至高维像素空间。然而,由于可行解空间受限,该飞行技能在探索中学习成本极高。为此,本文提出基于模型的轨迹优化器对智能体进行状态重置,以缓解此问题。所提出的训练流程与基线方法对比,并进行了消融实验以验证关键组件。下一步计划扩展缝隙尺寸与几何形状的变化,以激发涌现策略,并验证从仿真到现实的迁移能力。

原文摘要 · Abstract (English)

Flying through body-size narrow gaps in the environment is one of the most challenging moments for an underactuated multirotor. We explore a purely data-driven method to master this flight skill in simulation, where a neural network directly maps pixels and proprioception to continuous low-level control commands. This learned policy enables whole-body control through gaps with different geometries demanding sharp attitude changes (e.g., near-vertical roll angle). The policy is achieved by successive model-free reinforcement learning (RL) and online observation space distillation. The RL policy receives (virtual) point clouds of the gaps' edges for scalable simulation and is then distilled into the high-dimensional pixel space. However, this flight skill is fundamentally expensive to learn by exploring due to restricted feasible solution space. We propose to reset the agent as states on the trajectories by a model-based trajectory optimizer to alleviate this problem. The presented training pipeline is compared with baseline methods, and ablation studies are conducted to identify the key ingredients of our method. The immediate next step is to scale up the variation of gap sizes and geometries in anticipation of emergent policies and demonstrate the sim-to-real transformation.

无人机控制强化学习视觉导航端到端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。