用轨迹优化与强化学习结合,让太空机械臂更智能地完成维修任务。
Path Planning and Reinforcement Learning-Driven Control of On-Orbit Free-Flying Multi-Arm Robots
- 先用轨迹优化生成高效安全路径,再用强化学习自适应跟踪
- 仿真显示新方法在两种场景下均优于传统策略,动作更平稳
- 适合研究空间机器人自主控制的学者或航天工程人员
本文提出一种融合轨迹优化(TO)与强化学习(RL)的混合方法,用于在轨服务场景中自由漂浮多臂机器人的运动规划与控制。该系统利用TO生成满足动力学与运动学约束的可行高效路径,同时通过RL实现对不确定性的自适应轨迹跟踪。多臂机器人配备推进器,具备冗余性与稳定性,能有效减少对机械臂的稳定依赖,提升机动性。TO协同优化臂部运动与推进器力,增强操作效率;RL则通过无模型控制应对动态交互与干扰。通过全面仿真验证,两个案例研究——初始接触表面运动与需表面逼近的自由漂浮场景——均表明该方法优于传统策略。推进器显著提升了运动平滑性、安全性与效率,且RL策略能有效追踪TO生成轨迹,处理高维动作空间与动态不匹配问题。该融合框架结合了精准规划与强鲁棒性,为复杂动态空间环境下的机器人自主性提供坚实基础。
原文摘要 · Abstract (English)
This paper presents a hybrid approach that integrates trajectory optimization (TO) and reinforcement learning (RL) for motion planning and control of free-flying multi-arm robots in on-orbit servicing scenarios. The proposed system integrates TO for generating feasible, efficient paths while accounting for dynamic and kinematic constraints, and RL for adaptive trajectory tracking under uncertainties. The multi-arm robot design, equipped with thrusters for precise body control, enables redundancy and stability in complex space operations. TO optimizes arm motions and thruster forces, reducing reliance on the arms for stabilization and enhancing maneuverability. RL further refines this by leveraging model-free control to adapt to dynamic interactions and disturbances. The experimental results validated through comprehensive simulations demonstrate the effectiveness and robustness of the proposed hybrid approach. Two case studies are explored: surface motion with initial contact and a free-floating scenario requiring surface approximation. In both cases, the hybrid method outperforms traditional strategies. In particular, the thrusters notably enhance motion smoothness, safety, and operational efficiency. The RL policy effectively tracks TO-generated trajectories, handling high-dimensional action spaces and dynamic mismatches. This integration of TO and RL combines the strengths of precise, task-specific planning with robust adaptability, ensuring high performance in the uncertain and dynamic conditions characteristic of space environments. By addressing challenges such as motion coupling, environmental disturbances, and dynamic control requirements, this framework establishes a strong foundation for advancing the autonomy and effectiveness of space robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。