让机器人用少量示范学会新任务,还能持续进步不遗忘。
Few-Shot Vision-Language Action-Incremental Policy Learning
- 用多模态交互提取少数示范中的任务特征提示
- 通过任务关系图实现连续学习,成功率达基准26%以上
- 适合需要快速适应新任务的机器人系统研究者
近年来,基于Transformer的机器人操作方法利用多视角空间表示和语言指令,通过大量机器人示范学习运动轨迹。然而,机器人数据采集极为困难,现有方法缺乏在仅有少量示范情况下对新任务进行持续学习的能力。本文将此问题定义为少样本动作增量学习(FSAIL)任务,并设计了任务提示图演化策略(TOPIC)来应对挑战。为解决数据稀缺问题,TOPIC通过多模态信息深度交互,在少量示范中学习任务特定提示(TSP),有效提取任务特异性判别信息。为增强对新任务的持续学习能力并缓解灾难性遗忘,TOPIC采用连续演化策略(CES),利用任务间的内在关联构建任务关系图,从而通过复用先前任务习得技能,高效适应新任务。TOPIC首次在机器人操作任务中实现少样本持续学习,大量实验表明其成功率优于现有最先进基线模型超过26%,显著提升基于Transformer策略的持续学习能力。
原文摘要 · Abstract (English)
Recently, Transformer-based robotic manipulation methods utilize multi-view spatial representations and language instructions to learn robot motion trajectories by leveraging numerous robot demonstrations. However, the collection of robot data is extremely challenging, and existing methods lack the capability for continuous learning on new tasks with only a few demonstrations. In this paper, we formulate these challenges as the Few-Shot Action-Incremental Learning (FSAIL) task, and accordingly design a Task-prOmpt graPh evolutIon poliCy (TOPIC) to address these issues. Specifically, to address the data scarcity issue in robotic imitation learning, TOPIC learns Task-Specific Prompts (TSP) through the deep interaction of multi-modal information within few-shot demonstrations, thereby effectively extracting the task-specific discriminative information. On the other hand, to enhance the capability for continual learning on new tasks and mitigate the issue of catastrophic forgetting, TOPIC adopts a Continuous Evolution Strategy (CES). CES leverages the intrinsic relationships between tasks to construct a task relation graph, which effectively facilitates the adaptation of new tasks by reusing skills learned from previous tasks. TOPIC pioneers few-shot continual learning in the robotic manipulation task, and extensive experimental results demonstrate that TOPIC outperforms state-of-the-art baselines by over 26$\%$ in success rate, significantly enhancing the continual learning capabilities of existing Transformer-based policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。