统一可微框架提升四旋翼控制强化学习训练效果
VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control
- 构建跨任务的可微分四旋翼控制框架,支持悬停、跟踪等四大任务
- 提出ABPT算法,解决状态覆盖不足与梯度偏差问题,提升训练鲁棒性
- 验证了仿真策略在真实场景中的初步迁移能力,适合强化学习研究者
基于可微分仿真的第一阶强化学习在四旋翼控制中前景广阔,但实际进展分散于特定任务。为支持更系统的发展与评估,本文提出一个面向多任务四旋翼控制的统一可微分框架。该框架封装良好、可扩展,配备面向部署的动力学模型,为四个代表性任务——悬停、跟踪、着陆和竞速提供统一接口。同时,我们提出了首阶学习算法套件,识别出标准训练中的两大瓶颈:由时域初始化导致的状态覆盖受限,以及由部分不可微奖励引起的梯度偏差。为此,我们提出修正的反向传播通过时间(ABPT)方法,结合可微滚动优化、基于值函数的辅助目标和已访问状态初始化,显著提升训练稳定性。实验表明,ABPT在部分不可微奖励的任务中表现最优,同时在全可微设置下仍具竞争力。我们还提供了概念验证级的真实世界部署,展示了在该框架中学习的策略具备从仿真到现实的初步迁移能力。
原文摘要 · Abstract (English)
First-order reinforcement learning with differentiable simulation is promising for quadrotor control, but practical progress remains fragmented across task-specific settings. To support more systematic development and evaluation, we present a unified differentiable framework for multi-task quadrotor control. The framework is wrapped, extensible, and equipped with deployment-oriented dynamics, providing a common interface across four representative tasks: hovering, tracking, landing, and racing. We also present the suite of first-order learning algorithms, where we identify two practical bottlenecks of standard first-order training: limited state coverage caused by horizon initialization and gradient bias caused by partially non-differentiable rewards. To address these issues, we propose Amended Backpropagation Through Time (ABPT), which combines differentiable rollout optimization, a value-based auxiliary objective, and visited-state initialization to improve training robustness. Experimental results show that ABPT yields the clearest gains in tasks with partially non-differentiable rewards, while remaining competitive in fully differentiable settings. We further provide proof-of-concept real-world deployments showing initial transferability of policies learned in the proposed framework beyond simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。