arXiv:2410.14957cs.ROcs.AI2024-10被引 2

用少量示范实现图像抓取的持续优化,避免随机探索失败。

Offline-to-online Reinforcement Learning for Image-based Grasping with Scarce Demonstrations

  • 用神经正切核启发的正则化替代目标网络,稳定训练。
  • 仅需50次示范,2小时内成功率超90%。
  • 适合示范稀缺的现实机器人抓取任务,如真空抓取。

离线到在线强化学习(O2O RL)旨在让智能体在与环境交互过程中持续改进策略,同时保证初始策略行为可接受。这种可接受性对机器人操作至关重要,因随机探索可能导致严重失败和时间浪费。当仅有少量(可能次优)示范时,该场景下行为克隆(BC)易受分布偏移影响而失效。此前研究指出,在基于图像的环境中应用O2O RL面临挑战。本文提出一种新型O2O RL算法,可在真实图像驱动的机器人真空抓取任务中,仅用少量示范成功学习。该算法将离策略演员-评论家算法中的目标网络替换为受神经正切核启发的正则化技术。实验表明,仅需50次人类示范,该算法在两小时内即可达到超过90%的成功率,而传统BC及常用强化学习算法均无法达成类似性能。

原文摘要 · Abstract (English)

Offline-to-online reinforcement learning (O2O RL) aims to obtain a continually improving policy as it interacts with the environment, while ensuring the initial policy behaviour is satisficing. This satisficing behaviour is necessary for robotic manipulation where random exploration can be costly due to catastrophic failures and time. O2O RL is especially compelling when we can only obtain a scarce amount of (potentially suboptimal) demonstrations$\unicode{x2014}$a scenario where behavioural cloning (BC) is known to suffer from distribution shift. Previous works have outlined the challenges in applying O2O RL algorithms under the image-based environments. In this work, we propose a novel O2O RL algorithm that can learn in a real-life image-based robotic vacuum grasping task with a small number of demonstrations where BC fails majority of the time. The proposed algorithm replaces the target network in off-policy actor-critic algorithms with a regularization technique inspired by neural tangent kernel. We demonstrate that the proposed algorithm can reach above 90\% success rate in under two hours of interaction time, with only 50 human demonstrations, while BC and existing commonly-used RL algorithms fail to achieve similar performance.

强化学习机器人抓取少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。