用大模型预处理语言指令,提升机器人在复杂任务中的理解能力
IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- 用大视觉语言模型预处理指令,生成更优输入增强VLA
- 在含重复物体的任务中,性能显著提升,尤其面对新概念推理
- 适合研究复杂语义指令下的机器人操作与多模态模型优化
视觉-语言-动作模型(VLAs)近年来成为解决机器人操作问题的热门方法。然而,这些模型需以适合机器人控制的速率输出动作,限制了可使用的语言模型规模,从而影响其语言理解能力。操作任务可能需要复杂的语言指令,例如通过相对位置识别目标物体来表达人类意图。为此,我们提出IA-VLA框架,利用大型视觉语言模型的强语言理解能力作为预处理阶段,生成改进后的上下文以增强VLA的输入。我们在一组语义复杂的任务上评估该框架,这些任务在现有VLA文献中研究较少,即涉及视觉重复对象(视觉上无法区分的物体)的任务。使用包含三类场景的重复物体数据集,对比基线VLA与两种增强变体。实验表明,该增强方案使VLA受益,尤其是在需要从演示中推断新概念的语言指令下表现更优。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) have become an increasingly popular approach for addressing robot manipulation problems in recent years. However, such models need to output actions at a rate suitable for robot control, which limits the size of the language model they can be based on, and consequently, their language understanding capabilities. Manipulation tasks may require complex language instructions, such as identifying target objects by their relative positions, to specify human intention. Therefore, we introduce IA-VLA, a framework that utilizes the extensive language understanding of a large vision language model as a pre-processing stage to generate improved context to augment the input of a VLA. We evaluate the framework on a set of semantically complex tasks which have been underexplored in VLA literature, namely tasks involving visual duplicates, i.e., visually indistinguishable objects. A dataset of three types of scenes with duplicate objects is used to compare a baseline VLA against two augmented variants. The experiments show that the VLA benefits from the augmentation scheme, especially when faced with language instructions that require the VLA to extrapolate from concepts it has seen in the demonstrations. For the code, dataset, and videos, see https://sites.google.com/view/ia-vla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。