arXiv:2607.03461cs.CVcs.LG2026-07

用统一模型实现视觉语言动作世界建模,性能显著提升。

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

论文配图:WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
图 1 · 摘自论文原文
  • 基于BAGEL构建统一框架,整合视觉、语言、动作与环境建模。
  • 在LIBERO等数据集上超越专用模型,动作表征更结构化。
  • 适合研究多模态协同建模与机器人智能的学者参考。

世界模型旨在以支持感知、推理和行动的方式捕捉环境动态,已成为视觉-语言-动作-世界(VLAW)建模的核心方向。尽管统一视觉-语言模型展现出强大的多模态生成能力,其作为世界模型的潜力仍待探索。本文提出 exttt{WorldBagel},基于现代多模态统一模型BAGEL构建的统一VLAW框架,系统研究了统一架构在世界建模中的作用。在多任务机器人操作与跨域实验中, exttt{WorldBagel}持续优于专用模型,学习到更具结构性且与视觉和语言上下文语义对齐的动作表示。在LIBERO、Language Table和Franka数据集上的实验表明,统一不仅是架构便利,更是构建有效VLAW模型的关键因素,带来一致的实证收益,并深化对多模态世界建模的理解。代码与模型检查点将在接受后发布。

原文摘要 · Abstract (English)

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vision-language models have demonstrated strong multimodal generation capabilities, yet their potential as world models remains underexplored. In this work, we introduce \texttt{WorldBagel}, a unified VLAW framework built on BAGEL, a modern multimodal unified model, and use it to systematically investigate the role of unification in world modeling. Across multi-task robotic manipulation and cross-domain experiments, \texttt{WorldBagel} consistently outperforms task-specific alternatives and learns action representations that are more structured and semantically aligned with visual and linguistic context. Experiments on LIBERO, Language Table, and Franka show that unification is not only an architectural convenience, but also a key factor in learning effective VLAW models, leading to consistent empirical gains and deeper insights into multimodal world modeling. Code and model checkpoints will be released upon acceptance.

多模态建模机器人学习统一模型世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。