用强化学习+动态估计算法提升无人机机械臂在不确定环境下的操作精度。
Reinforcement Learning with Inner-loop Dynamics Estimator for Aerial Manipulation under Uncertainty

- 分层控制:上层强化学习规划动作,下层动态估计补偿扰动。
- 硬件实验显示末端误差更小,任务成功率显著提升。
- 适合需高鲁棒性空中操作的机器人研究者参考。
空中机械臂可实现难以到达区域的物理交互;然而,在快速臂运动、负载变化及未知动态不确定性条件下,直接进行整体机身控制的问题仍基本未解。本文提出一种分层控制框架,结合强化学习(RL)与内环动态估计算法以应对该挑战。上层强化学习将期望的六自由度末端执行器目标映射为协调的全身指令,实现无需精确耦合动力学模型的任务驱动控制。下层则通过无需系统模型知识的动态估计方案,实时跟踪指令并补偿瞬时惯性变化与不确定性。我们在自研四旋翼搭载三自由度机械臂的硬件平台上,于不同负载条件下验证了该方法。相比RL+PID与RL+INDI+PID基线,所提方法显著降低末端跟踪误差,提升任务成功率。结果表明,将学习得来的全身协同与基于估计器的低层补偿相结合,可有效提升空中操作在工况变化下的精度与鲁棒性。
原文摘要 · Abstract (English)
Aerial manipulators enable physical interaction in hard-to-reach environments; however, the combined problem of direct whole-body aerial manipulation under rapid arm motion, payload changes, and related unknown dynamic uncertainty remains a largely unsolved problem. We present a hierarchical control framework that combines Reinforcement Learning (RL) with an inner-loop dynamics estimator to address this problem. The RL outer loop maps desired 6-degrees-of-freedom (DOF) end-effector targets to coordinated whole-body commands, enabling direct task-driven control without relying on a fully accurate coupled dynamic model in the policy layer. An inner loop then tracks these commands while compensating for transient inertial shifts and uncertainty during execution via a dynamics estimator scheme without requiring system model knowledge. We validate the proposed approach on a custom quadrotor equipped with a 3-DoF manipulator through hardware experiments under varying payload conditions. Compared with RL+PID and RL+INDI+PID baselines, the proposed method reduces end-effector tracking error and improves task success rate across the tested hardware conditions. These results show that combining learned whole-body coordination with estimator-based low-level compensation improves the precision and robustness of aerial manipulation under changing operating conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。