EO-1让机器人在开放世界中实现视觉、语言与动作的无缝协同控制。
EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- 统一架构处理图像、文本、视频和动作,支持多模态交互。
- 基于150万条高质量数据训练,显著提升复杂任务泛化能力。
- 适合研究通用机器人控制与多模态决策的开发者参考。
人类在开放世界中无缝进行多模态推理与物理交互的能力,是通用具身智能系统的核心目标。近年来,视觉-语言-动作(VLA)模型通过大规模机器人与视觉文本数据联合训练,在通用机器人控制方面取得显著进展,但仍难以达到人类水平的交错推理与交互灵活性。本文提出EO-Robotics,包含EO-1模型与EO-Data1.5M数据集。EO-1是一个统一的具身基础模型,通过交错式视觉-文本-动作预训练,在多模态具身推理与机器人控制上表现卓越。其构建基于两大支柱:(i) 统一架构可无差别处理图像、文本、视频和动作;(ii) 大规模高质量多模态具身推理数据集EO-Data1.5M,包含超过150万样本,强调视觉-文本-动作的交错理解。EO-1在EO-Data1.5M上通过自回归解码与流匹配去噪的协同训练,实现无缝机器人动作生成与多模态具身推理。大量实验验证了交错式视觉-文本-动作学习在开放世界理解与泛化中的有效性,涵盖多种形态的长时序精细操作任务。本文详细阐述了EO-1的架构、EO-Data1.5M的数据构建策略及训练方法,为先进具身基础模型的发展提供宝贵洞见。
原文摘要 · Abstract (English)
The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on large-scale robot and visual-text data, have demonstrated notable progress in general robot control. However, they still fail to achieve human-level flexibility in interleaved reasoning and interaction. In this work, we introduce EO-Robotics, consists of EO-1 model and EO-Data1.5M dataset. EO-1 is a unified embodied foundation model that achieves superior performance in multimodal embodied reasoning and robot control through interleaved vision-text-action pre-training. The development of EO-1 is based on two key pillars: (i) a unified architecture that processes multimodal inputs indiscriminately (image, text, video, and action), and (ii) a massive, high-quality multimodal embodied reasoning dataset, EO-Data1.5M, which contains over 1.5 million samples with emphasis on interleaved vision-text-action comprehension. EO-1 is trained through synergies between auto-regressive decoding and flow matching denoising on EO-Data1.5M, enabling seamless robot action generation and multimodal embodied reasoning. Extensive experiments demonstrate the effectiveness of interleaved vision-text-action learning for open-world understanding and generalization, validated through a variety of long-horizon, dexterous manipulation tasks across multiple embodiments. This paper details the architecture of EO-1, the data construction strategy of EO-Data1.5M, and the training methodology, offering valuable insights for developing advanced embodied foundation models. Project Page: https://eo-robotics.ai/eo-1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。