用自然语言指挥机器人完成复杂任务,还能自动补救技能不足。
CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

- 用少量演示学习技能,结合预训练视觉语言模型理解语义
- 执行时能组合技能、绑定参数,成功率最高达100%
- 发现能力缺口会主动要新示范,无需重新训练
让机器人理解并执行自然语言指令,同时保持数据高效性仍是挑战。视觉-语言-动作(VLA)和视觉-语言模型(VLMs)提供直观交互方式,但需要大量数据;任务参数化模仿学习虽数据高效,却缺乏自然语言基础。本文通过模块化架构,将任务参数化核化运动基元(TP-KMPs)与预训练的VLM结合,实现优势互补。学习阶段,仅需2至5次动力学示范,由VLM生成描述每个技能参数与先决条件的技能模板。执行阶段,VLM解析指令,选择合适技能,推理参数绑定,并通过协方差加权组合生成新行为。当无既定技能或组合可用时,系统可识别能力缺口并请求针对性示范,全程无需微调。在7自由度机械臂上的验证表明,在涉及技能选择、组合与主动学习的场景中,成功率范围为73.3%至100%。
原文摘要 · Abstract (English)
Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation models such as vision-language-action (VLA) and vision-language models (VLMs) provide intuitive interaction channels but require extensive data; task-parameterized imitation learning achieves data efficiency but lacks natural language grounding. This work bridges this gap through a modular architecture combining task-parameterized kernelized movement primitives (TP-KMPs) with pretrained VLMs. During learning, skills are acquired from 2 to 5 kinesthetic demonstrations, and the VLM generates skill schemas describing each skill's parameters and preconditions. During execution, the VLM interprets commands to select skills, reason about parameter bindings, and create novel behaviors through covariance-weighted composition. When no skill or composition suffices, the system identifies capability gaps and requests targeted demonstrations, all without fine-tuning. Validation on a 7-DoF manipulator shows success rates of 73.3%-100% in scenarios requiring skill selection, composition, and active learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。