arXiv:2605.30877cs.RO2026-05被引 14

40亿参数视觉语言动作模型可直接控制机器人,无需微调就完成多项任务。

Wall-OSS-0.5 Technical Report

论文配图:Wall-OSS-0.5 Technical Report
图 1 · 摘自论文原文
  • 用多模态协同训练让大模型直接生成机器人动作,不依赖下游微调。
  • 零样本在17个任务上达成可观进展,微调后平均任务进度达60.5%。
  • 保留视觉语言理解能力的同时增强具身智能,适合具身智能研究者。

大规模视觉-语言-动作(VLA)预训练正成为机器人策略的基础,但现有证据几乎都来自特定任务微调后的表现。这留下一个根本问题:VLA预训练本身是否能产生可执行的机器人行为,还是仅提供更好的下游初始化?我们提出Wall-OSS-0.5,一个开源的40亿参数VLA模型,基于30亿参数的视觉语言模型(VLM)骨干,增加动作生成组件,旨在使预训练的机器人能力可在真实硬件上直接测量。模型在超过20种机器人形态上进行预训练,每轮处理超过一百万条机器人轨迹,并结合一个接地的多模态语料库。采用梯度桥接联合训练方法,三个目标各司其职:离散动作预测将强的VLM原生梯度引入骨干网络,多模态预测保持接地的视觉-语言理解,连续流匹配作为部署时的动作接口。在未进行任务微调的情况下,预训练检查点已实现非平凡的零样本真实机器人行为,在17个任务套件中完成多个任务,包括一个未见过的可变形物体操作任务,取得高任务进展。微调后,同一检查点作为更强的适应先验,在15个真实机器人任务上达到60.5%的平均任务进度,优于π_0.5模型17.5%。多模态评估进一步证实,动作训练未损害视觉-语言能力:模型在强化具身接地的同时,仍保持广泛的视觉-语言能力。这些结果将VLA预训练从一种初始化策略重新定位为可直接测试、已有实用价值的机器人能力来源。

原文摘要 · Abstract (English)

Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs is almost invariably reported after task-specific fine-tuning. This leaves a foundational question unanswered: does VLA pretraining itself yield executable robot behavior, or does it merely furnish a better initialization for downstream policy learning? We present Wall-OSS-0.5, an open-source 4B VLA built upon a 3B VLM backbone augmented with action-generation components, designed so that pretrained robotic capability is directly measurable on physical hardware. The model is pretrained across more than 20 embodiments, processing over one million robot trajectories per epoch alongside a grounded multimodal corpus. We adopt a gradient-bridged co-training recipe in which three objectives play distinct and complementary roles: discrete action prediction routes strong VLM-native gradients into the backbone, multimodal prediction preserves grounded vision-language understanding, and continuous flow matching serves as the deployment-time action interface. Before task-specific fine-tuning, the pretrained checkpoint achieves non-trivial zero-shot real-robot behavior, completing several tasks, including a held-out deformable manipulation task, at high task progress on a 17-task suite. After fine-tuning, the same checkpoint serves as a stronger adaptation prior, reaching 60.5% average task progress on 15 real-robot tasks and outperforming π_0.5 by 17.5%. Multimodal evaluations further confirm that action training does not erode grounded vision-language competence: the model preserves broad vision-language ability while strengthening embodied grounding. Together, these results reposition VLA pretraining from an initialization strategy to a directly testable, already useful source of robot capability.

机器人具身智能多模态零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。