arXiv:2510.18337cs.RO2025-10被引 9

提出统一快慢推理的视觉语言动作模型,提升机器人指令响应速度与准确性。

MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning

  • 采用混合变压器架构,融合通用视觉语言模型与领域专家网络。
  • 在仿真与真实机器人上实现更快推理和更强语言可控性。
  • 适合需要高效多任务决策的智能机器人系统研究者使用。

将视觉语言指令融入视觉运动策略正成为提升机器人开放世界泛化能力的重要方向。现有方法面临两大挑战:缺乏生成推理作为条件时语言控制力不足,或引入推理导致显著延迟。本文提出MoTVLA,一种基于混合变压器(MoT)的视觉-语言-动作(VLA)模型,集成快慢统一推理与行为策略学习。MoTVLA保留预训练视觉语言模型(作为通用智能体)在感知、场景理解与语义规划等任务上的通用能力,同时引入一个共享知识的领域专家(第二层Transformer),生成特定领域的快速推理(如机器人运动分解),从而提升策略执行效率。通过以分解后的运动指令为条件,该模型可学习多样化行为,显著增强语言可控性。在自然语言处理基准、机器人仿真环境及真实世界实验中的广泛评估表明,MoTVLA在快慢推理与操作任务性能上均表现更优。

原文摘要 · Abstract (English)

Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existing approaches face two challenges: limited language steerability when no generated reasoning is used as a condition, or significant inference latency when reasoning is incorporated. In this work, we introduce MoTVLA, a mixture-of-transformers (MoT)-based vision-language-action (VLA) model that integrates fast-slow unified reasoning with behavior policy learning. MoTVLA preserves the general intelligence of pre-trained VLMs (serving as the generalist) for tasks such as perception, scene understanding, and semantic planning, while incorporating a domain expert, a second transformer that shares knowledge with the pretrained VLM, to generate domain-specific fast reasoning (e.g., robot motion decomposition), thereby improving policy execution efficiency. By conditioning the action expert on decomposed motion instructions, MoTVLA can learn diverse behaviors and substantially improve language steerability. Extensive evaluations across natural language processing benchmarks, robotic simulation environments, and real-world experiments confirm the superiority of MoTVLA in both fast-slow reasoning and manipulation task performance.

机器人学习视觉语言推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。