用强化学习让无人机在狭窄缝隙中高精度、高动态飞行。
Precise Aggressive Aerial Maneuvers with Sensorimotor Policies
- 直接从机载视觉和本体感知映射到低层控制指令
- 实现5厘米间隙、90度倾斜的高精度穿越,重复性好
- 无需预先知道缝隙位置,可应对移动缝隙,适合复杂环境
轻量级机载传感器下实现精准高动态无人机机动仍是关键瓶颈。此类机动对拓展系统可访问区域至关重要,尤其在通过环境中的狭窄开口时。代表性问题为四旋翼在SE(3)约束下穿越狭窄缝隙,需利用短暂倾斜姿态和机体非对称性完成。本文通过强化学习(RL)在仿真中端到端训练传感器-运动策略,将机载视觉与本体感知直接映射为低层控制命令。采用基于模型规划器生成轨迹的初始化策略,缓解模型无关强化学习在受限解空间中的探索难题。精心设计的仿真到现实迁移使策略能稳定控制四旋翼通过低间隙(如5厘米)且高重复性地穿越。例如,方法可在未知缝隙位置与朝向情况下,实现高达90度倾斜的矩形缝隙穿越。未在动态缝隙上训练的情况下,策略仍可实时响应并穿越移动缝隙。该方法还在紧密排列的多个狭窄缝隙挑战赛道上验证有效。策略学习方法的灵活性还体现在无需人工定义穿越姿态或视觉特征,即可适应几何多样的缝隙。
原文摘要 · Abstract (English)
Precise aggressive maneuvers with lightweight onboard sensors remains a key bottleneck in fully exploiting the maneuverability of drones. Such maneuvers are critical for expanding the systems' accessible area by navigating through narrow openings in the environment. Among the most relevant problems, a representative one is aggressive traversal through narrow gaps with quadrotors under SE(3) constraints, which require the quadrotors to leverage a momentary tilted attitude and the asymmetry of the airframe to navigate through gaps. In this paper, we achieve such maneuvers by developing sensorimotor policies directly mapping onboard vision and proprioception into low-level control commands. The policies are trained using reinforcement learning (RL) with end-to-end policy distillation in simulation. We mitigate the fundamental hardness of model-free RL's exploration on the restricted solution space with an initialization strategy leveraging trajectories generated by a model-based planner. Careful sim-to-real design allows the policy to control a quadrotor through narrow gaps with low clearances and high repeatability. For instance, the proposed method enables a quadrotor to navigate a rectangular gap at a 5 cm clearance, tilted at up to 90-degree orientation, without knowledge of the gap's position or orientation. Without training on dynamic gaps, the policy can reactively servo the quadrotor to traverse through a moving gap. The proposed method is also validated by training and deploying policies on challenging tracks of narrow gaps placed closely. The flexibility of the policy learning method is demonstrated by developing policies for geometrically diverse gaps, without relying on manually defined traversal poses and visual features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。