用分层强化学习实现0.1mm精度的圆柱装配,兼顾稳定与灵活。
Knowledge-Guided Hierarchical Policy Learning for High-Precision Cylindrical Assembly under Tight Tolerances

- 下层融合专家经验与TD3算法,上层用规则动态调整决策
- 500次内收敛,极端条件下仍保持高成功率
- 适合精密装配场景,尤其对误差容忍度要求高的工业应用
提出一种混合分层学习框架,实现170mm圆柱组件在0.1mm公差下的高精度装配。底层网络通过行为克隆(BC)融入专家经验,赋予机器人类人直觉,并结合双延迟深度确定性策略梯度(TD3)提升训练稳定性与鲁棒性;上层网络基于启发式规则动态调整底层决策,保障操作灵活性。构建仿真模型进行预训练,实现高效安全的现实迁移。实验表明,奖励曲线在500个回合内完成收敛,展现出高效学习能力;对初始条件和位姿误差具有更强适应性,在极端条件下仍能保持良好成功率;同时在高斯噪声干扰下表现出优异稳定性。真实世界中,圆柱段装配轨迹更加平滑,波动更小。
原文摘要 · Abstract (English)
A hybrid hierarchical learning framework is proposed to achieve high-precision assembly of 170mm cylindrical components with tolerance of 0.1mm. The lower-level network integrates expert experience through Behavior Cloning (BC), giving the robot human-like intuition, and incorporates the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to enhance training stability and robustness. The upper-level network dynamically adjusts the lower-level decisions based on heuristic rules, ensuring flexibility in operations. A simulated model is constructed to learn before transferring to real world. An efficient and safe training is allowed. Comparisons show that the reward curve converges within 500 episodes, indicating high learning efficiency. It also demonstrates better adaptability to initial conditions and pose errors, achieving satisfactory success rates even under extreme conditions. Moreover, the method exhibits good stability under Gaussian noise interference. In the real world, the assembly trajectory of the cylindrical segment shows smoother motion and less fluctuation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。