让微型飞艇在真实环境中稳定倒挂,突破传统控制极限。
Learning Robust Control Policies for Inverted Pose on Miniature Blimp Robots
- 用高保真仿真+改进TD3算法训练倒挂控制策略
- 仿真中成功率高于传统能量整形控制器
- 通过映射层实现仿真到真实环境的无缝部署
实现并维持倒挂姿态对发挥微型飞艇机器人(MBRs)的全部机动性至关重要。然而,由于其复杂的欠驱动动力学,开发可靠的倒挂控制策略仍具挑战。为此,我们提出一种新框架,实现MBR倒挂姿态的鲁棒控制策略学习。该框架包含三个核心阶段:首先,基于真实MBR运动数据构建并校准高保真三维(3D)仿真环境;其次,在仿真中使用改进的双延迟深度确定性策略梯度(TD3)算法结合领域随机化策略训练鲁棒倒挂控制策略;第三,设计映射层以弥合仿真到现实的差距,促进所学策略在真实环境中的部署。仿真环境中的综合评估表明,所学策略的成功率高于能量整形控制器。此外,实验结果证实,加入映射层的策略可使MBR在真实场景中实现并维持完全倒挂姿态。
原文摘要 · Abstract (English)
The ability to achieve and maintain inverted poses is essential for unlocking the full agility of miniature blimp robots (MBRs). However, developing reliable inverted control strategies for MBRs remains challenging due to their complex and underactuated dynamics. To address this challenge, we propose a novel framework that enables robust control policy learning for inverted pose on MBRs. The proposed framework consists of three core stages. First, a high-fidelity three-dimensional (3D) simulation environment is constructed and calibrated using real-world MBR motion data. Second, a robust inverted control policy is trained in simulation using a modified Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm combined with a domain randomization strategy. Third, a mapping layer is designed to bridge the sim-to-real gap and facilitate real-world deployment of the learned policy. Comprehensive evaluations in the simulation environment demonstrate that the learned policy achieves a higher success rate compared to the energy-shaping controller. Furthermore, experimental results confirm that the learned policy with a mapping layer enables an MBR to achieve and maintain a fully inverted pose in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。