机器人通过对话持续学习新技能,用少量示范即可掌握新任务。
Continual Robot Skill and Task Learning via Dialogue
- 用对话向人类询问未知技能,结合低秩适配的视觉运动策略
- 在模拟中对新技能提升超300%,旧技能表现相当
- 真实用户实验中成功教机器人做饭,且分心任务完成率更高
交互式机器人学习面临挑战:机器人需在人机互动中高效持续学习新技能以应对新任务。本文提出一个框架,使机器人通过与人类用户的对话交互,持续学习新任务和视觉-运动技能,并从真人处查询未知技能。机器人维护一个技能库,利用现有大语言模型进行具身对话,获取新技能知识。我们设计了新的视觉-运动控制策略ACT-LoRA,仅需少量示范即可持续学习新技能,这对人机交互场景至关重要。论文有两个目标:一是在仿真中展示更优的持续学习性能;二是在真实人机交互场景中验证对话学习框架的有效性。结果表明,ACT-LoRA在多个持续学习基准上显著优于GMM-LoRA基线,新技能性能提升超过300%,旧技能表现持平。此外,经伦理审批的人体实验显示,该对话框架可100%成功教会用户烹饪技能,且测试阶段用户用于完成辅助分心任务的时间比例显著高于非学习型语言代理(p < 0.001)。
原文摘要 · Abstract (English)
Interactive robot learning is a challenging problem as the robot is present with human users who expect the robot to learn novel skills to solve novel tasks perpetually with sample efficiency. In this work we present a framework for robots to continually learn tasks and visuo-motor skills and query for novel skills via dialog interactions with human users. Our robot agent maintains a skill library, and uses an existing LLM to perform grounded dialog interactions to query unknown skills from real human users. We developed a novel visual-motor control policy Action Chunking Transformer with Low Rank Adaptation (ACT-LoRA) that can continually learn novel skills using only a few demonstrations which is critical in human-robot interaction scenarios. The paper has twin goals: Firstly to demonstrate better continual learning in simulation; and secondly, to demonstrate the use of our dialog based learning framework in a realistic human-robot interaction use case. Our ACT-LoRA policy consistently outperforms a GMM-LoRA baseline on multiple continual learning simulation benchmarks by achieving > 300% improvements on novel skills, while achieving comparable performance in existing skills. Moreover, with our IRB approved human-subjects study we demonstrate that our dialog based continual learning framework allows users to teach robots cooking skills successfully (100%) while spending a higher ratio of time on finishing an auxiliary distraction tasks in the test phase of the study compared to a non-learning language based agent (p < 0.001).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。