arXiv:2410.18519cs.ROcs.SY2024-10被引 4

用学习环境训练软体机器人控制器,无需先验知识即可实现高效闭环控制。

Reinforcement Learning Controllers for Soft Robots using Learned Environments

  • 基于数据生成合成环境,用策略梯度方法训练控制器。
  • 在长时序任务中实现高性能闭环控制,无需了解机器人特性。
  • 适合希望快速部署无模型控制的软体机器人研究者。

软体机械臂因其柔性和可变形结构具有操作优势,但其固有的非线性动力学带来巨大挑战。传统解析方法常依赖简化假设,而基于学习的方法计算开销大且受限于已有数据。本文提出一种新方法,利用先进的策略梯度算法,在从数据中学习得到的可并行化合成环境中进行软体机器人控制。我们设计了一种安全导向的执行空间探索协议,通过级联更新与加权随机性实现。具体地,采用递归前向动力学模型,通过物理上安全的均值回归随机游走生成训练数据,以探索部分可观测的状态空间。实验表明,基于先进演员-评论家方法的强化学习框架能高效学习长时序的高性能行为。该方法无需任何关于机器人运行或能力的知识,为软体机器人控制提供全面基准工具。

原文摘要 · Abstract (English)

Soft robotic manipulators offer operational advantage due to their compliant and deformable structures. However, their inherently nonlinear dynamics presents substantial challenges. Traditional analytical methods often depend on simplifying assumptions, while learning-based techniques can be computationally demanding and limit the control policies to existing data. This paper introduces a novel approach to soft robotic control, leveraging state-of-the-art policy gradient methods within parallelizable synthetic environments learned from data. We also propose a safety oriented actuation space exploration protocol via cascaded updates and weighted randomness. Specifically, our recurrent forward dynamics model is learned by generating a training dataset from a physically safe \textit{mean reverting} random walk in actuation space to explore the partially-observed state-space. We demonstrate a reinforcement learning approach towards closed-loop control through state-of-the-art actor-critic methods, which efficiently learn high-performance behaviour over long horizons. This approach removes the need for any knowledge regarding the robot's operation or capabilities and sets the stage for a comprehensive benchmarking tool in soft robotics control.

强化学习软体机器人仿真训练无模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。