让视觉语言模型学会在真实空间中行动,突破动作理解瓶颈。
Igniting VLMs toward the Embodied Space
- 用端到端架构融合指令推理与动作生成,实现跨层级思维链统一。
- 在长程操作任务中成功率显著超越基线,能精准执行复杂指令。
- 适合研究具身智能、机器人控制或多模态大模型落地的开发者。
尽管基础模型在语言和视觉领域取得显著进展,现有视觉语言模型(VLMs)仍缺乏对空间与具身性的理解。将VLMs迁移至具身领域暴露出模态、预训练分布与训练目标之间的根本性不匹配,导致动作理解与生成成为通往通用人工智能(AGI)的核心瓶颈。我们提出WALL-OSS,一种端到端的具身基础模型,通过大规模多模态预训练实现(1)具身感知的视觉-语言理解,(2)强语言-动作关联,以及(3)稳健的操作能力。该方法采用紧密耦合架构与多策略训练课程,使统一跨层级思维链(Unified Cross-Level CoT)无缝整合指令推理、子目标分解与细粒度动作合成于单一可微框架中。实验表明,WALL-OSS在复杂长程操作任务中表现优异,具备强大指令遵循能力、复杂理解与推理能力,并显著优于多个强基线模型,为从VLMs迈向具身基础模型提供了一条可靠且可扩展的路径。
原文摘要 · Abstract (English)
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferring VLMs to embodied domains reveals fundamental mismatches between modalities, pretraining distributions, and training objectives, leaving action comprehension and generation as a central bottleneck on the path to AGI. We introduce WALL-OSS, an end-to-end embodied foundation model that leverages large-scale multimodal pretraining to achieve (1) embodiment-aware vision-language understanding, (2) strong language-action association, and (3) robust manipulation capability. Our approach employs a tightly coupled architecture and multi-strategies training curriculum that enables Unified Cross-Level CoT-seamlessly unifying instruction reasoning, subgoal decomposition, and fine-grained action synthesis within a single differentiable framework. Our results show that WALL-OSS attains high success on complex long-horizon manipulations, demonstrates strong instruction-following capabilities, complex understanding and reasoning, and outperforms strong baselines, thereby providing a reliable and scalable path from VLMs to embodied foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。