用图文交错指令提升机器人零样本泛化能力,支持手绘图等灵活输入。
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
- 采用图文交错输入,让机器人理解更自然的任务指令
- 在未见物体上泛化能力提升2倍,零样本支持手绘图等多样输入
- 自动生成大规模真实世界图文指令数据集,适合研究人机交互与通用机器人
基础模型的兴起为物理世界中的通用机器人策略铺平了道路。现有依赖纯文本指令的方法在面对未见场景时往往表现不佳。我们认为,图文交错输入能提供更丰富、更少偏倚的上下文,使机器人能更好地处理未见任务并实现更灵活的人机交互。基于此,我们提出Interleave-VLA,首个能理解图文交错指令并直接生成连续动作序列的机器人学习范式。该方法无需改变现有视觉-语言-动作(VLA)模型架构,仅需微小调整即可实现强零样本泛化。此外,我们构建了一个自动管道,将Open X-Embodiment中的文本指令转化为图文交错形式,生成包含210万条轨迹的大规模真实世界多模态数据集。仿真与真实世界评估表明,Interleave-VLA具有两大优势:(1) 对未见物体的域外泛化能力比纯文本基线提升2倍;(2) 可零样本支持灵活任务接口和多样化指令,如手绘草图。我们归因于其使用指令图像有效缓解幻觉,并融合来自互联网的异构多模态数据,具备可扩展潜力。更多信息请访问 https://interleave-vla.github.io/Interleave-VLA-Anonymous/
原文摘要 · Abstract (English)
The rise of foundation models paves the way for generalist robot policies in the physical world. Existing methods relying on text-only instructions often struggle to generalize to unseen scenarios. We argue that interleaved image-text inputs offer richer and less biased context and enable robots to better handle unseen tasks with more versatile human-robot interaction. Building on this insight, Interleave-VLA, the first robot learning paradigm capable of comprehending interleaved image-text instructions and directly generating continuous action sequences in the physical world, is introduced. It offers a natural, flexible, and model-agnostic paradigm that extends state-of-the-art vision-language-action (VLA) models with minimal modifications while achieving strong zero-shot generalization. Interleave-VLA also includes an automatic pipeline that converts text instructions from Open X-Embodiment into interleaved image-text instructions, resulting in a large-scale real-world interleaved embodied dataset with 210k episodes. Comprehensive evaluation in simulation and the real world shows that Interleave-VLA offers two major benefits: (1) improves out-of-domain generalization to unseen objects by 2x compared to text input baselines, (2) supports flexible task interfaces and diverse instructions in a zero-shot manner, such as hand-drawn sketches. We attribute Interleave-VLA's strong zero-shot capability to the use of instruction images, which effectively mitigate hallucinations, and the inclusion of heterogeneous multimodal datasets, enriched with Internet-sourced images, offering potential for scalability. More information is available at https://interleave-vla.github.io/Interleave-VLA-Anonymous/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。