arXiv:2602.12062cs.RO2026-02被引 11

HoloBrain-0让机器人具备更强3D空间理解能力,支持低成本部署。

HoloBrain-0 Technical Report

  • 融合机器人结构与视觉语言的新型架构,提升三维推理能力。
  • 0.2B参数版本在仿真和真实任务中表现媲美更大模型。
  • 开源完整工具链,支持从数据到部署的全流程复现。

本文提出HoloBrain-0,一种融合视觉、语言与动作的综合框架,弥合基础模型研究与真实机器人部署之间的差距。系统核心为新型视觉-语言-动作(VLA)架构,显式引入多视角相机参数与运动学描述(URDF),增强3D空间推理并支持多种机器人形态。通过可扩展的“预训练后微调”范式,在RoboTwin 2.0、LIBERO和GenieSim等仿真基准上取得领先性能,并在复杂长周期真实操作任务中表现优异。其高效的0.2B参数变体性能媲美显著更大的基线模型,实现低延迟设备端部署。为加速研究与应用,我们开源整个HoloBrain生态:(1)强大的预训练VLA基础模型;(2)多套仿真与真实任务的微调检查点;(3)RoboOrchard——覆盖数据采集、模型训练与部署的全栈VLA基础设施。结合标准化数据采集协议,本发布为高性能机器人操作提供完整、可复现的技术路径。

原文摘要 · Abstract (English)

In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot deployment. The core of our system is a novel VLA architecture that explicitly incorporates robot embodiment priors, including multi-view camera parameters and kinematic descriptions (URDF), to enhance 3D spatial reasoning and support diverse embodiments. We validate this design through a scalable ``pre-train then post-train" paradigm, achieving state-of-the-art results on simulation benchmarks such as RoboTwin 2.0, LIBERO, and GenieSim, as well as strong results on challenging long-horizon real-world manipulation tasks. Notably, our efficient 0.2B-parameter variant rivals significantly larger baselines, enabling low-latency on-device deployment. To further accelerate research and practical adoption, we fully open-source the entire HoloBrain ecosystem, which includes: (1) powerful pre-trained VLA foundations; (2) post-trained checkpoints for multiple simulation suites and real-world tasks; and (3) RoboOrchard, a full-stack VLA infrastructure for data curation, model training and deployment. Together with standardized data collection protocols, this release provides the community with a complete, reproducible path toward high-performance robotic manipulation.

机器人视觉语言开源部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。