用固定非线性动态网络提升机器人模仿学习的稳定性
Learning from Demonstration with Implicit Nonlinear Dynamics Models
- 引入可调非线性动力学的循环层建模时间序列
- 在手写任务中误差累积减少,精度显著提升
- 适合需要长期稳定输出的机器人控制场景
模仿学习(LfD)是训练复杂运动任务策略的有效方法,如机器人操作。实际应用中,需克服执行过程中误差积累导致的漂移问题,即错误随时间累积引发分布外行为。现有方法通过扩大数据采集、人工干预修正、时间集成预测或学习具有收敛保证的动力系统模型来应对。本文提出一种新方法:受储备计算启发,设计一个包含可调非线性动力学特性的固定动态系统的循环神经网络层,用于建模时序动态。我们在LASA人类手写数据集上验证该层有效性,实证表明将其嵌入现有网络架构可有效缓解模仿学习中的误差累积问题。与时间集成预测和回声状态网络(ESN)相比,本方法在手写任务中表现出更高精度与鲁棒性,并能泛化至多种动力学环境,同时保持较低延迟。
原文摘要 · Abstract (English)
Learning from Demonstration (LfD) is a useful paradigm for training policies that solve tasks involving complex motions, such as those encountered in robotic manipulation. In practice, the successful application of LfD requires overcoming error accumulation during policy execution, i.e. the problem of drift due to errors compounding over time and the consequent out-of-distribution behaviours. Existing works seek to address this problem through scaling data collection, correcting policy errors with a human-in-the-loop, temporally ensembling policy predictions or through learning a dynamical system model with convergence guarantees. In this work, we propose and validate an alternative approach to overcoming this issue. Inspired by reservoir computing, we develop a recurrent neural network layer that includes a fixed nonlinear dynamical system with tunable dynamical properties for modelling temporal dynamics. We validate the efficacy of our neural network layer on the task of reproducing human handwriting motions using the LASA Human Handwriting Dataset. Through empirical experiments we demonstrate that incorporating our layer into existing neural network architectures addresses the issue of compounding errors in LfD. Furthermore, we perform a comparative evaluation against existing approaches including a temporal ensemble of policy predictions and an Echo State Network (ESN) implementation. We find that our approach yields greater policy precision and robustness on the handwriting task while also generalising to multiple dynamics regimes and maintaining competitive latency scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。