用强化学习让太空机械臂安全抓取碎片,不撞自己也不碰目标。
Safe Obstacle-Free Guidance of Space Manipulators in Debris Removal Missions via Deep Reinforcement Learning
- 用双评价网络分工:一个盯准目标,一个防碰撞。
- 在模拟七自由度机械臂上实现无故障轨迹规划,成功率高。
- 适合做空间碎片清理的机器人控制研究者参考。
本研究旨在通过基于Twin Delayed Deep Deterministic Policy Gradient(TD3)的模型无关工作空间轨迹规划器,实现空间机械臂在碎片清除任务中的安全可靠捕获。采用具有奇异性规避与可操作性增强的局部控制策略,确保执行稳定。机械臂需同时追踪非合作目标上的捕获点、避免自碰撞,并防止与目标发生意外接触。为此,提出一种基于课程学习的多评价网络架构,其中一个评价器强调精准跟踪,另一个强化碰撞规避。同时使用优先级经验回放缓冲区以加速收敛并提升策略鲁棒性。该框架在Matlab/Simulink中基于安装于自由浮动基座上的七自由度KUKA LBR iiwa机械臂进行仿真验证,证明了其在碎片清除任务中生成安全且自适应轨迹的有效性。
原文摘要 · Abstract (English)
The objective of this study is to develop a model-free workspace trajectory planner for space manipulators using a Twin Delayed Deep Deterministic Policy Gradient (TD3) agent to enable safe and reliable debris capture. A local control strategy with singularity avoidance and manipulability enhancement is employed to ensure stable execution. The manipulator must simultaneously track a capture point on a non-cooperative target, avoid self-collisions, and prevent unintended contact with the target. To address these challenges, we propose a curriculum-based multi-critic network where one critic emphasizes accurate tracking and the other enforces collision avoidance. A prioritized experience replay buffer is also used to accelerate convergence and improve policy robustness. The framework is evaluated on a simulated seven-degree-of-freedom KUKA LBR iiwa mounted on a free-floating base in Matlab/Simulink, demonstrating safe and adaptive trajectory generation for debris removal missions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。