arXiv:2605.11697cs.RO2026-05

用强化学习实现双机器人精密插孔,提升成功率与稳定性。

Rainbow Deep Q-Learning with Kinematics-Aware Design for Cooperative Delta and 3-RRS Parallel Robot Insertion

论文配图:Rainbow Deep Q-Learning with Kinematics-Aware Design for Cooperative Delta and 3-RRS Parallel Robot Insertion
图 1 · 摘自论文原文
  • 先优化3-RRS机械臂几何结构,扩大安全操作空间。
  • 采用彩虹DQN框架,12维状态+12种动作,实现稳定插孔。
  • 适合工业装配、多机器人协同控制场景,性能优于传统方法。

本文提出一种基于彩虹深度Q网络(Rainbow DQN)的运动学感知强化学习框架,用于协作式Delta并联机器人与3-RRS(转动-转动-球副)并联机械臂完成插孔任务。核心贡献在于学习前引入几何设计优化阶段:通过调整3-RRS结构以最大化无奇异性工作空间并改善条件性,从而扩大强化学习策略可探索的安全区域。两机械臂共同实现6自由度可控子空间(3个Delta平移、2个3-RRS旋转、1个3-RRS垂直移动);由于插孔任务对销轴方向旋转不变,任务相关流形为5维。该协作插入问题被建模为马尔可夫决策过程,状态向量为12维,动作集包含12个离散增量指令(每个受控自由度正负各一)。奖励函数结合密集邻近引导、运动学与工作空间违规惩罚,以及成功插入的稀疏奖励。彩虹DQN融合双重Q学习、斗篷架构、优先经验回放、多步回报、噪声线性层探索和分布值头,在两阶段课程下训练。在高保真运动学仿真中验证,该协同设计框架实现了策略稳定收敛、可靠插入,并显著降低约束违反,优于原始DQN代理和经典采样规划器。

原文摘要 · Abstract (English)

This paper presents a kinematics-aware deep reinforcement learning framework based on Rainbow Deep Q-Networks (DQN) for cooperative peg-in-hole manipulation by a Delta parallel robot and a 3-RRS (Revolute--Revolute--Spherical) parallel manipulator. A key contribution is the integration of a geometric design-optimization stage that precedes learning: the 3-RRS geometry is tuned to maximize the singularity-free workspace and improve conditioning, which in turn enlarges the safe region in which the reinforcement learning policy can explore. Together the two manipulators expose a 6~degree-of-freedom (DoF) controllable subspace (three Delta translations, two 3-RRS rotations, and one 3-RRS vertical translation); the peg-in-hole task is invariant to rotation about the peg axis, so the task-relevant manifold is five dimensional. The cooperative insertion problem is cast as a Markov Decision Process with a 12-dimensional state vector and a discrete action set containing $6 \times 2 = 12$ incremental commands (one positive and one negative per controlled DoF). A shaped reward combines dense proximity guidance, penalties for kinematic and workspace violations, and sparse bonuses for successful insertions. The Rainbow DQN -- integrating double Q-learning, dueling architecture, prioritized replay, multi-step returns, noisy linear layers for exploration, and a distributional value head -- is trained with a two-stage curriculum. The co-designed framework is validated in a high-fidelity kinematic simulator, where it achieves stable policy convergence, reliable insertions, and reduced constraint violations compared against a vanilla DQN agent and a classical sampling-based planner.

强化学习机器人控制插孔装配多机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。