用大模型重标注机器人数据,让机械臂更听懂指令。
Task Robustness via Re-Labelling Vision-Action Robot Data

- 用视觉语言模型分解任务为带语义的子步骤
- 生成含物体属性的多样指令,提升泛化能力
- 适合想提升机器人指令理解力的研究者
当前机器人学习中大规模模型虽能完成多种操作并适应新场景,但仍难以准确执行指令,主要因现有数据集的语言和动作序列多样性不足。本文提出TREAD框架,利用预训练视觉语言模型(VLM)在不额外采集数据的前提下增强已有机器人数据集,通过三阶段流程:从原始指令和场景生成语义子任务,基于子任务分割演示视频,生成包含物体属性的多样化指令,从而将长示范分解为有语义基础的语言-动作对。同时通过增加语言多样性的目标文本进一步提升鲁棒性。在LIBERO数据集上的评估表明,使用增强数据训练的策略在未见任务和目标上表现更优,证明TREAD在轨迹分解和语言条件策略泛化方面均有显著提升。
原文摘要 · Abstract (English)
The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios. However, these policies continue to struggle with following instructions, likely due to the limited linguistic and action sequence diversity in existing robotics datasets. This paper introduces Task Robustness via Re-Labelling Vision-Action Robot Data (TREAD), a scalable framework that leverages large Vision-Language Models (VLMs) to augment existing robotics datasets without additional data collection, harnessing the transferable knowledge embedded in these models. Our approach leverages a pretrained VLM through three stages: generating semantic sub-tasks from original instruction labels and initial scenes, segmenting demonstration videos conditioned on these sub-tasks, and producing diverse instructions that incorporate object properties, effectively decomposing longer demonstrations into grounded language-action pairs. We further enhance robustness by augmenting the data with linguistically diverse versions of the text goals. Evaluations on LIBERO demonstrate that policies trained on our augmented datasets exhibit improved performance on novel, unseen tasks and goals. Our results show that TREAD enhances both planning generalization through trajectory decomposition and language-conditioned policy generalization through increased linguistic diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。