arXiv:2608.07546cs.ROcs.AI2026-08

用电机级策略让机器人跨配置通用,训练一次就能控制不同数量电机的缆绳机器人。

Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies

论文配图:Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies
图 1 · 摘自论文原文
  • 在电机层面控制每根缆绳长度,而非直接控制末端位置。
  • 1个策略可适配任意电机数量和布局的缆绳机器人,2D仿真策略成功控制8电机3D实机。
  • 无需逆运动学,避免复杂正运动学问题,适合快速部署到新机器人结构。

缆绳驱动并联机器人(CDPRs)具有多种构型和复杂的控制挑战,可通过深度强化学习(DRL)学习其非线性动力学来解决。然而,现有DRL方法通常需要大量训练时间,且所学策略难以泛化到不同机器人构型或不同数量的执行器。本文提出一种新型DRL方法,通过电机级策略控制每个电机以达到目标缆绳长度,而非传统方法中直接控制机器人末端到达目标位置。据我们所知,这是首个使用电机级策略控制CDPRs的DRL工作。该方法具备两大优势:(i) 单一共享策略可适用于任意构型的CDPR,不受执行器数量影响;(ii) 避免了更复杂的正运动学问题,依赖逆运动学求解。训练在仿真中完成,所学策略成功迁移到真实机器人。实验结果表明,电机级策略(ALP)在鲁棒性和精度上均优于传统强化学习方法。我们进一步用一个在2D平面训练的4电机仿真策略,成功控制了一个实际的8电机3D CDPR实现三维运动,验证了该方法对任意构型、任意电机数量和布局的适用性。

原文摘要 · Abstract (English)

Cable-driven parallel robots (CDPRs) present diverse configurations and complex control challenges, which can be addressed by deep reinforcement learning (DRL) by learning their nonlinear dynamics. However, DRL methods often require extensive training time, and the resulting policies do not generalize well to different robot configurations or varying numbers of actuators. In this article, we introduce a novel DRL approach for controlling CDPRs that does not depend on the specific robot configuration. Our method trains an actuator-level policy that controls each motor to achieve its target cable length, in contrast to conventional DRL approaches that learn to control the entire robot to reach a desired end-effector position. To the best of our knowledge, this is the first work to apply DRL to control CDPRs using an actuator-level policy. This approach offers two main advantages: (i) a single shared policy can be applied to any CDPR configuration, regardless of actuator count, and (ii) reliance on inverse kinematics, avoiding the more challenging forward kinematics problem. Training is performed in simulation, and the learned policy is successfully transferred to a real CDPR. Experimental results show that the actuator-level policy (ALP) surpasses traditional reinforcement learning methods in both robustness and precision. We further control a real 8-motor CDPR with 3D motion using a policy trained on a simulated 4-motor planar CDPR operating in 2D. This illustrates that the proposed method is applicable to any CDPR configuration, independent of actuator number or placement.

强化学习机器人控制泛化能力缆绳机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。