Orca通过统一世界表征,实现对世界的预测与行动。
Orca: The World is in Your Mind

- 以状态转移为核心,融合视频与语言学习世界模型
- 125K小时视频+160M事件标注,构建大规模训练数据
- 通用世界表征支持文本、图像、动作生成,适合多模态研究
我们提出Orca,首个通用世界基础模型。它从多模态世界信号中学习统一的世界潜在空间,并通过多模态读出接口暴露该空间。不同于孤立的下一个词、下一帧或下一步预测,我们聚焦于下一状态预测建模,提供统一的状态转移路径以理解、预测和作用于世界。Orca通过两种互补范式学习:无意识学习从连续视频中捕捉密集自然状态转移,有意识学习通过语言描述事件和VQA监督建模稀疏有意义的状态转移。预训练阶段,我们构建大规模世界学习数据集,包含125,000小时视频数据和1600万事件标注。预训练后,Orca学习到统一的世界潜在空间。为检验其下游能力,我们评估三种代表性读出任务:文本生成、图像预测和具身动作生成。模型主干冻结,仅可训练轻量级模态特定解码器。实验表明该范式具有可扩展性,更强的世界潜在空间带来更强的下游读出表现。Orca在同等规模下优于专用基线模型。结果表明,Orca作为通用世界基础模型,为理解、预测和作用于世界提供了有前景的方法。最后,我们讨论当前局限,旨在为社区提供启示。
原文摘要 · Abstract (English)
We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multimodal readout interfaces. Rather than optimizing isolated next-token, next-frame, or next-action prediction, we are centered on Next-State-Prediction modeling, offering a unified state-transition modeling route toward understanding, predicting, and acting upon the world. Orca learns through two complementary paradigms: unconscious learning captures dense natural state transitions from continuous videos, and conscious learning models sparse meaningful state transitions by language-described events and VQA supervision. For pre-training, we construct a large-scale world-learning inventory data, including 125K hours of video data and 160M event annotations. After pre-training, Orca learns a unified world latent space. To examine whether the learned latent supports downstream, we evaluate it by three representative downstream readouts: text generation, image prediction, and embodied action generation. Orca's backbone is frozen, and only the lightweight modality-specific decoders are trainable. Experiments show the scalability of the proposed paradigm and verify that stronger world latent enables stronger downstream readouts. Orca outperforms similar-sized specialized baselines. These results show that Orca, as a general world foundation model, presents a promising approach to understanding, predicting, and acting upon the world. Finally, we discuss the current limitations, aiming to provide useful insights and inspiration for the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。