arXiv:2510.11027cs.CV2025-10被引 23

让机器人同时理解语言和动作,实现更聪明的自主决策。

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

  • 构建融合高层推理与底层控制的视觉-语言-动作模型
  • 在多个机器人任务中达到当前最优表现,如空间推理和任务规划
  • 揭示预训练模型对机器人控制微调的影响,适合研究具身智能的学者

尽管已有大量研究致力于利用视觉-语言模型(VLM)发展具身推理能力,或将先进VLM集成到视觉-语言-动作(VLA)模型中以实现端到端机器人控制,但很少有工作直接解决上游基于VLM的推理与下游VLA策略学习之间的关键差距。本文提出Vlaser——一种具备协同具身推理能力的视觉-语言-动作模型,是一个面向具身代理的基线视觉-语言模型,旨在将高层推理与底层控制相融合。基于高质量的Vlaser-6M数据集,Vlaser在多个具身推理基准测试中表现优异,涵盖空间推理、具身定位、具身问答和任务规划等。此外,我们系统分析了不同VLM初始化对监督式VLA微调的影响,为缓解互联网规模预训练数据与具身特定策略学习数据之间的领域偏移提供了新见解。基于此,该方法在WidowX基准上取得当前最优结果,在Google Robot基准上也展现出竞争力。

原文摘要 · Abstract (English)

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies directly address the critical gap between upstream VLM-based reasoning and downstream VLA policy learning. In this work, we take an initial step toward bridging embodied reasoning with VLA policy learning by introducing Vlaser - a Vision-Language-Action Model with synergistic embodied reasoning capability, which is a foundational vision-language model designed to integrate high-level reasoning with low-level control for embodied agents. Built upon the high-quality Vlaser-6M dataset, Vlaser achieves state-of-the-art performance across a range of embodied reasoning benchmarks - including spatial reasoning, embodied grounding, embodied QA, and task planning. Furthermore, we systematically examine how different VLM initializations affect supervised VLA fine-tuning, offering novel insights into mitigating the domain shift between internet-scale pre-training data and embodied-specific policy learning data. Based on these insights, our approach achieves state-of-the-art results on the WidowX benchmark and competitive performance on the Google Robot benchmark.

具身智能视觉语言机器人控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。