arXiv:2409.02636cs.ROcs.SY2024-09被引 15

用Mamba提升机器人模仿学习的运动编码能力

Mamba as a motion encoder for robotic imitation learning

  • 将Mamba作为运动编码器,压缩序列信息并保留时序动态
  • 在杯具摆放和箱子装载任务中成功率达92%以上,优于Transformer
  • 适合数据有限场景下的实时运动生成应用

近期模仿学习结合大语言模型技术的进展,显著提升了机器人的灵巧性与适应性。本文提出将Mamba——一种前沿架构,已在大语言模型中展现潜力——用于机器人模仿学习,强调其作为编码器有效捕捉上下文信息的能力。通过降低状态空间维度,Mamba以类似自编码器的方式运作,能将序列信息高效压缩为状态变量,同时保持运动预测所需的关键时序动态。在杯具放置与箱体装载等任务上的实验表明,尽管估计误差略高,但Mamba在实际任务执行中的成功率显著优于Transformer,这一表现归因于其状态空间模型结构。研究还验证了Mamba在少量训练数据下作为实时运动生成器的可行性。

原文摘要 · Abstract (English)

Recent advancements in imitation learning, particularly with the integration of LLM techniques, are set to significantly improve robots' dexterity and adaptability. This paper proposes using Mamba, a state-of-the-art architecture with potential applications in LLMs, for robotic imitation learning, highlighting its ability to function as an encoder that effectively captures contextual information. By reducing the dimensionality of the state space, Mamba operates similarly to an autoencoder. It effectively compresses the sequential information into state variables while preserving the essential temporal dynamics necessary for accurate motion prediction. Experimental results in tasks such as cup placing and case loading demonstrate that despite exhibiting higher estimation errors, Mamba achieves superior success rates compared to Transformers in practical task execution. This performance is attributed to Mamba's structure, which encompasses the state space model. Additionally, the study investigates Mamba's capacity to serve as a real-time motion generator with a limited amount of training data.

机器人学习Mamba运动编码模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。