用高保真可微仿真实现四旋翼端到端控制,飞行更稳更快。
Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation

- 直接从深度图输出体速率指令,跳过传统路径规划和外环控制器
- 在多个基准上成功率最高,最大速度达7.5m/s, jerk 最低
- 无需复杂网络或预训练,零样本部署到未见户外环境
障碍物避让是四旋翼执行高级任务的基础视觉任务。现有优化与学习方法通常将四旋翼建模为质点,生成路径或速度指令后由外环控制器跟踪,但在高速飞行时,规划轨迹常超出控制器能力,导致动态不可行。本文提出一种新型端到端策略,通过强化学习结合可微仿真,直接从深度图像映射到低层体速率命令。训练中经参数辨识的高保真仿真显著缩小了训练、仿真与真实世界间的差距。可微仿真提供精确梯度,确保无专家指导下高效训练低层策略。该策略采用轻量、最简推理流程,无需显式映射、主干网络、基础结构、递归模块或后端控制器,也无需课程学习或特权信息。通过直接向硬件控制器输出低层指令,实现全飞行包线控制,避免动态不可行问题。实验表明,该方法在多个基准上取得最高成功率与最低 jerk,且具备强泛化能力,在未见过的室外环境中实现零样本部署,最高速度达7.5m/s,可在超密集森林中稳定飞行。
原文摘要 · Abstract (English)
Obstacle avoidance is a fundamental vision-based task essential for enabling quadrotors to perform advanced applications. When planning the trajectory, existing approaches both on optimization and learning typically regard quadrotor as a point-mass model, giving path or velocity commands then tracking the commands by outer-loop controller. However, at high speeds, planned trajectories sometimes become dynamically infeasible in actual flight, which beyond the capacity of controller. In this paper, we propose a novel end-to-end policy that directly maps depth images to low-level bodyrate commands by reinforcement learning via differentiable simulation. The high-fidelity simulation in training after parameter identification significantly reduces all the gaps between training, simulation and real world. Analytical process by differentiable simulation provides accurate gradient to ensure efficiently training the low-level policy without expert guidance. The policy employs a lightweight and the most simple inference pipeline that runs without explicit mapping, backbone networks, primitives, recurrent structures, or backend controllers, nor curriculum or privileged guidance. By inferring low-level command directly to the hardware controller, the method enables full flight envelope control and avoids the dynamic-infeasible issue.Experimental results demonstrate that the proposed approach achieves the highest success rate and the lowest jerk among state-of-the-art baselines across multiple benchmarks. The policy also exhibits strong generalization, successfully deploying zero-shot in unseen, outdoor environments while reaching speeds of up to 7.5m/s as well as stably flying in the super-dense forest. This work is released at https://github.com/Fanxing-LI/avoidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。