一个模型统一理解、推理、想象和行动,性能不妥协。
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

- 用同一视觉语言模型统一处理感知、推理与未来预测
- 单次前向传播生成任务链、动作与未来视频,性能领先
- 适合研究多模态智能体的开发者和研究人员
我们提出 Pelican-Unify 1.0,首个遵循统一原则训练的具身基础模型。该模型使用单一视觉语言模型(VLM)作为统一理解模块,将场景、指令、视觉上下文与动作历史映射到共享语义空间。同一 VLM 还作为统一推理模块,在单次前向传播中自回归生成面向任务、动作和未来的思维链,并将最终隐藏状态投影为密集潜在变量。一个统一未来生成器(UFG)基于该潜在变量,在同一去噪过程中通过两个模态特定输出头联合生成未来视频与动作。语言、视频与动作损失均反向传播至共享表示,实现理解、推理、想象与行动的联合优化,而非训练三个独立专家系统。实验表明,统一不等于妥协:单个检查点在八项 VLM 基准上达 64.7 分,同类模型中最佳;在 WorldArena 上得 66.03 分,排名第一;在 RoboTwin 上达 93.5 分,动作类方法中排名第二。结果证明,统一范式在保持专业能力的同时,成功融合了理解、推理、想象与行动。
原文摘要 · Abstract (English)
We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding module, mapping scenes, instructions, visual contexts, and action histories into a shared semantic space. The same VLM also serves as a unified reasoning module, autoregressively producing task-, action-, and future-oriented chains of thought in a single forward pass and projecting the final hidden state into a dense latent variable. A Unified Future Generator (UFG) then conditions on this latent variable and jointly generates future videos and future actions through two modality-specific output heads within the same denoising process. The language, video, and action losses are all backpropagated into the shared representation, enabling the model to jointly optimize understanding, reasoning, imagination, and action during training, rather than training three isolated expert systems. Experiments demonstrate that unification does not imply compromise. With a single checkpoint, Pelican-Unify 1.0 achieves strong performance across all three capabilities: 64.7 on eight VLM benchmarks, the best among comparable-scale models; 66.03 on WorldArena, ranking first; and 93.5 on RoboTwin, the second-best average among compared action methods. These results show that the unified paradigm succeeds in preserving specialist strength while bringing understanding, reasoning, imagination, and action into one model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。