arXiv:2411.10991cs.ROcs.AI2024-11

用强化学习动态调节随机神经网络,让机器人学会新动作而无需重新训练。

Modulating Reservoir Dynamics via Reinforcement Learning for Efficient Robot Skill Synthesis

  • 用强化学习生成上下文信号,实时调控随机神经网络输出
  • 在2自由度机器人上实现无须重训练的障碍物避让与目标达成功能
  • 适合需要快速扩展动作库的机器人演示学习场景

随机循环神经网络(称为储层)可基于编码任务目标的上下文输入,学习机器人运动。通过线性回归将上下文调制的储层动态映射到期望轨迹,该储层计算(RC)方法无需迭代梯度下降,计算高效。本文提出一种新型基于储层计算的从示范学习(LfD)框架,不仅可学习示范动作,还能在线调节储层动态,生成初始示范集未覆盖的运动轨迹。这通过一个强化学习(RL)模块实现,该模块根据机器人状态学习输出上下文作为动作。由于上下文维度通常较低,此RL学习效率高。我们在一个2自由度模拟机器人上系统验证了该模型:机器人被教授以不同目标(上下文编码)进行抓取,且具备或不具避障约束。初始数据集包含一组抓取示范,由储层系统学习。为实现分布外目标的可达性,激活RL模块学习生成动态上下文,使生成轨迹达成目标,且无需对储层系统进行任何学习。整体模型利用初始学习到的动作基元集,结合设计的奖励函数,高效生成多样化运动行为,因此可作为灵活高效的从示范学习系统,可在不收集新数据的情况下扩展动作库。

原文摘要 · Abstract (English)

A random recurrent neural network, called a reservoir, can be used to learn robot movements conditioned on context inputs that encode task goals. The Learning is achieved by mapping the random dynamics of the reservoir modulated by context to desired trajectories via linear regression. This makes the reservoir computing (RC) approach computationally efficient as no iterative gradient descent learning is needed. In this work, we propose a novel RC-based Learning from Demonstration (LfD) framework that not only learns to generate the demonstrated movements but also allows online modulation of the reservoir dynamics to generate movement trajectories that are not covered by the initial demonstration set. This is made possible by using a Reinforcement Learning (RL) module that learns a policy to output context as its actions based on the robot state. Considering that the context dimension is typically low, learning with the RL module is very efficient. We show the validity of the proposed model with systematic experiments on a 2 degrees-of-freedom (DOF) simulated robot that is taught to reach targets, encoded as context, with and without obstacle avoidance constraint. The initial data set includes a set of reaching demonstrations which are learned by the reservoir system. To enable reaching out-of-distribution targets, the RL module is engaged in learning a policy to generate dynamic contexts so that the generated trajectory achieves the desired goal without any learning in the reservoir system. Overall, the proposed model uses an initial learned motor primitive set to efficiently generate diverse motor behaviors guided by the designed reward function. Thus the model can be used as a flexible and effective LfD system where the action repertoire can be extended without new data collection.

机器人学习强化学习储层计算动作泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。