用反事实标签增强数据,让机器人更懂复杂指令。
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- 用视觉语言模型生成反事实标签,丰富数据语义多样性。
- 无需新数据采集,指令跟随成功率翻倍。
- 适合需要精准理解复杂指令的机器人研究者。
通用机器人需能理解并执行用户指令。尽管视觉-语言-动作(VLA)模型可将开放词汇语言指令映射为机器人动作,但现有模型在执行细粒度指令时仍表现不佳。主要原因在于现有机器人数据集缺乏语义多样性与语言对齐,尤其是相似观测下的细粒度任务差异不足。为此,本文提出一种新方法:利用视觉语言模型生成反事实标签,对现有数据集进行增强。通过引入这些标签,提升了机器人数据集的语言对齐精度与任务粒度,从而显著改善VLA模型的指令遵循能力。我们在3个室内外环境中开展视觉-语言导航实验,评估模型从简单物体中心指令到复杂指代任务的执行能力。结果表明,反事实重标注(无需额外数据采集)显著提升VLA策略的指令跟随性能,优于当前最优方法,成功率较未增强数据训练的VLA模型提高一倍。此外,在包含干扰物的操控任务中,该方法同样带来明显性能提升。
原文摘要 · Abstract (English)
Generalist robots should be able to understand and follow user instructions. Despite providing a powerful architecture for mapping open-vocabulary language instructions to robot actions, current vision-language-action (VLA) models struggle to follow fine-grained commands. One cause for this is a lack of semantic diversity and language grounding in existing robot datasets and, specifically, a lack of fine-grained task diversity for similar observations. To address this, we present a novel method to augment existing robot datasets by leveraging vision-language models to create counterfactual labels. By augmenting existing datasets with these labels, we increase the diversity and granularity of language grounding for robot datasets, ultimately improving the language-following capabilities of VLAs. We evaluate the resulting model's ability to follow language instructions, ranging from simple object-centric commands to complex referential tasks, by conducting vision-language navigation experiments in 3 different indoor and outdoor environments. Our experiments show that counterfactual relabeling (without additional data collection) significantly improves instruction-following in VLA policies, outperforming state-of-the-art methods and doubling the success rate compared to VLAs trained on unaugmented data. We also evaluate our method for manipulation VLAs and find a similar gain in performance on tasks with distractors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。