arXiv:2606.01241cs.RO2026-06被引 1

一个模型同时搞定导航和操作,让机器人更通用。

OneVLA: A Unified Framework for Embodied Tasks

论文配图:OneVLA: A Unified Framework for Embodied Tasks
图 1 · 摘自论文原文
  • 统一架构设计,用同一动作头处理导航与操作任务。
  • 多阶段训练提升双任务协同,实测性能超越现有模型。
  • 适合想构建通用机器人的研究者与开发者。

导航与操作是具身智能的核心能力,使机器人能够理解自然语言指令并与其环境进行物理交互。然而,当前视觉-语言-动作(VLA)模型受限于任务专用架构,仅擅长导航或操作,阻碍了通用机器人代理的发展。为此,我们提出OneVLA,一种将两类任务整合到单一框架中的统一架构。具体而言,设计了一个可生成导航与操作动作的统一动作头,无需为不同任务定制。此外,提出多阶段渐进式训练策略,结合精心构建的数据与思维链(CoT)微调,促进两任务间的强正向迁移与相互增强。在模拟与真实环境中的大量实验表明,OneVLA达到领先性能,显著优于单任务专用及现有跨任务模型。通过统一核心能力,OneVLA为真正通用的机器人系统铺平道路。模型与源码将公开发布。

原文摘要 · Abstract (English)

Navigation and manipulation are fundamental capabilities of embodied intelligence, enabling robots to interpret natural language commands and interact physically with their surroundings. However, current Vision-Language-Action (VLA) models remain constrained by task-specific architectures, specializing in either navigation or manipulation, which hinders the development of general-purpose robotic agents. To bridge this gap, we introduce OneVLA, a unified architecture that integrates these distinct tasks into a single, cohesive framework. Specifically, we design a unified action head capable of generating both navigation and manipulation actions without requiring task-specific variants. Furthermore, we propose a multi stage progressive training strategy-incorporating curated data construction and Chain-of-Thought (CoT) fine-tuning that facilitates strong positive transfer and mutual reinforcement between the two domains. Extensive experiments in both simulated and real-world environments demonstrate that OneVLA achieves state-of-the-art performance, significantly outperforming both specialized single-task and existing cross-task models. By unifying these core capabilities, OneVLA paves the way for truly general-purpose robotic systems. The model and source code will be publicly released.

具身智能多任务学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。