用双层强化学习让欠驱动飞艇通过移动重心精准追踪目标
Bi-Level Reinforcement Learning Control for an Underactuated Blimp via Center-of-Mass Reconfiguration

- 分层设计:外层定重心位置,内层调推力跟踪路径
- 27个目标测试中精度与鲁棒性显著优于传统方法
- 适合对能效和载重有要求的轻型飞行器控制场景
本文研究通过重心重构实现欠驱动飞艇的目标导向轨迹跟踪。与依赖冗余推进的常规过驱动设计不同,本文采用仅含两个推进器和一个可移动内部滑块的紧凑结构,旨在提升能效和载荷能力。该硬件高效配置引入强非线性耦合与显著欠驱动特性。为此,提出一种双层强化学习框架,显式解耦任务级重心规划与连续推力控制。外层策略在飞行前确定目标相关的重心配置,内层策略生成推力指令以跟踪直线参考轨迹。为确保稳定学习,设计两阶段训练策略,并提供相应双层过程的收敛性分析。在包含27个目标的仿真与真实实验中,所提方法持续优于固定重心基线与基于PID的控制器,实现更高跟踪精度、更强鲁棒性及可靠的模拟到现实迁移性能。
原文摘要 · Abstract (English)
This paper investigates goal-directed tracking control of underactuated blimps with center-of-mass (CoM) reconfiguration. Unlike conventional overactuated blimp designs that rely on redundant actuation for simplified control, this paper focuses on a compact architecture consisting of two thrusters and a movable internal slider, aiming to improve energy efficiency and payload capacity. This hardware-efficient configuration introduces significant underactuation and strong nonlinear coupling between CoM dynamics and vehicle motion. To address these challenges, this paper proposes a bi-level reinforcement learning framework that explicitly decouples task-level CoM planning from continuous thrust control. The outer policy determines a target-dependent CoM configuration prior to flight, while the inner policy generates thrust commands to track straight-line references. To ensure stable learning, this paper introduces a two-stage learning strategy, supported by a convergence analysis of the resulting bi-level process. Extensive simulations and real-world experiments on a 27-goal evaluation set demonstrate that the proposed method consistently outperforms fixed-CoM baselines and PID-based controllers, achieving higher tracking accuracy, enhanced robustness, and reliable sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。