用强化学习让四轴无人机从任意姿态自恢复,仅靠简单传感器
Agile Fall Recovery for Quadrotors with Bidirectional Thrust via Reinforcement Learning

- 用循环策略+不对称架构应对传感器不完整和失灵问题
- 真实场景下零样本迁移,风扰和负载变化下均能成功恢复
- 无需状态估计,仅靠有限传感器实现敏捷自恢复
自主跌倒恢复是四轴无人机在真实环境中运行的关键能力,因碰撞或故障可能导致机体以任意姿态静止于地面。该问题挑战性高,因恢复需在有限机载感知、受限自由空间、存在地面接触及未知干扰条件下完成。本文提出一种基于强化学习的框架,仅使用轻量级机载传感器,实现四轴无人机从任意地面姿态恢复至稳定悬停。为应对严重部分可观测性和间歇性传感器失效,采用异构演员-评论家架构训练循环策略,并结合增量非线性动态逆(INDI)控制器跟踪策略输出。结合高保真电机响应与视觉光流仿真,整体训练框架显著缩小了仿真到现实的差距。仿真消融实验验证了主要设计选择的重要性,真实世界实验展示了零样本迁移能力,在不同初始姿态、风扰及附加负载下均实现鲁棒恢复。结果表明,仅依赖有限且不可靠的机载传感,即可实现敏捷四轴无人机跌倒恢复。
原文摘要 · Abstract (English)
Autonomous fall recovery is a critical capability for quadrotors operating in real-world environments, where collisions or failures may leave the vehicle resting on the ground in an arbitrary attitude. This problem is challenging because recovery must be achieved under limited onboard sensing, in constrained free space, with ground contact, and in the presence of unknown disturbances. In this letter, we present an RL-based framework for autonomous fall recovery of a quadrotor from arbitrary ground attitudes to stable hover using only lightweight onboard sensors. To address severe partial observability and intermittent sensor invalidity, we train a recurrent policy within an asymmetric actor--critic architecture, leveraging an Incremental Nonlinear Dynamic Inversion (INDI) controller to track the policy output. Combined with high-fidelity simulations of motor response and optical flow, the overall training framework significantly reduces the sim-to-real gap. Simulation ablation studies validate the importance of the main design choices, while real-world experiments demonstrate zero-shot transfer and robust recovery under different initial attitudes, wind disturbances, and additional payloads. These results demonstrate that agile quadrotor fall recovery can be achieved without explicit state estimation using only limited and unreliable onboard sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。