让大模型学会组合低级动作完成复杂机器人任务
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
- 构建基于ManiSkill2的机器人推理基准ClevrSkills
- 模型在组合任务上表现差,即使预训练数据量巨大
- 适合研究多模态大模型与机器人组合推理的学者
机器人任务本质上具有高度组合性。例如,执行'清理桌面'这类高层任务,需依次完成移动机械臂、抓取物体并移开等底层操作,同时动态评估环境变化。鉴于大型视觉语言模型(VLM)在需要人类级推理的任务上取得进展,我们提出问题:若模型掌握必要底层能力,能否自主组合这些能力以完成新颖的高层任务(如清理桌面),而无需显式教学?为此,我们提出了ClevrSkills——一个面向机器人组合推理的基准套件。该套件基于ManiSkill2模拟器开发,包含带语言和视觉标注的轨迹数据集,以及多模态提示作为任务规范。套件包含三个层级的组合理解任务,从基础运动技能开始。我们在ClevrSkills上对多个VLM基线进行评测,结果表明,即使经过大量任务的预训练,这些模型仍无法有效处理机器人任务中的组合推理。
原文摘要 · Abstract (English)
Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effectors to the objects on the table, pick them up and then move them off the table one-by-one, while re-evaluating the consequently dynamic scenario in the process. Given that large vision language models (VLMs) have shown progress on many tasks that require high level, human-like reasoning, we ask the question: if the models are taught the requisite low-level capabilities, can they compose them in novel ways to achieve interesting high-level tasks like cleaning the table without having to be explicitly taught so? To this end, we present ClevrSkills - a benchmark suite for compositional reasoning in robotics. ClevrSkills is an environment suite developed on top of the ManiSkill2 simulator and an accompanying dataset. The dataset contains trajectories generated on a range of robotics tasks with language and visual annotations as well as multi-modal prompts as task specification. The suite includes a curriculum of tasks with three levels of compositional understanding, starting with simple tasks requiring basic motor skills. We benchmark multiple different VLM baselines on ClevrSkills and show that even after being pre-trained on large numbers of tasks, these models fail on compositional reasoning in robotics tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。