统一视觉语言动作模型,能自适应推理与执行复杂任务
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning

- 单模型融合推理与动作,关键时切换至显式推理
- 在长序列任务中实现高精度规划与错误恢复
- 适合需要自然交互与复杂操作的通用机器人场景
通用机器人需具备多样任务的协同推理与执行能力。现有双系统方法将高层推理与底层动作分离,常面临系统间能力理解不足和延迟问题。本文提出OneTwoVLA,一个统一的视觉-语言-动作模型,兼具动作生成(系统一)与推理(系统二)能力。其核心在于:在任务执行的关键时刻主动进行显式推理,其余时间则基于最近一次推理结果生成动作。为激发模型的推理与泛化能力,我们设计了一套可扩展的具身推理导向视觉语言数据合成流水线,用于与机器人数据联合训练。通过大量实验验证,OneTwoVLA在四大关键能力上表现卓越:长时程任务规划、错误检测与恢复、自然人机交互、可泛化的视觉定位,使其能够完成如制作火锅、调制鸡尾酒等长序列高灵巧度操作任务。
原文摘要 · Abstract (English)
General-purpose robots capable of performing diverse tasks require synergistic reasoning and acting capabilities. However, recent dual-system approaches, which separate high-level reasoning from low-level acting, often suffer from challenges such as limited mutual understanding of capabilities between systems and latency issues. This paper introduces OneTwoVLA, a single unified vision-language-action model that can perform both acting (System One) and reasoning (System Two). Crucially, OneTwoVLA adaptively switches between two modes: explicitly reasoning at critical moments during task execution, and generating actions based on the most recent reasoning at other times. To further unlock OneTwoVLA's reasoning and generalization capabilities, we design a scalable pipeline for synthesizing embodied reasoning-centric vision-language data, used for co-training with robot data. We validate OneTwoVLA's effectiveness through extensive experiments, highlighting its superior performance across four key capabilities: long-horizon task planning, error detection and recovery, natural human-robot interaction, and generalizable visual grounding, enabling the model to perform long-horizon, highly dexterous manipulation tasks such as making hotpot or mixing cocktails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。