arXiv:2606.02800cs.CVcs.AI2026-06被引 61

Cosmos 3统一建模语言、图像、视频、音频与动作,打造通用物理智能引擎。

Cosmos 3: Omnimodal World Models for Physical AI

论文配图:Cosmos 3: Omnimodal World Models for Physical AI
图 1 · 摘自论文原文
  • 采用混合Transformer架构,支持多模态输入输出灵活组合。
  • 在图文生成、图像转视频等任务上达开源模型最佳水平。
  • 适合构建具身智能体,推动物理AI开放研究与部署。

我们提出Cosmos 3,一类面向物理智能的多模态世界模型,可在统一的混合Transformer架构中联合处理与生成语言、图像、视频、音频及动作序列。通过支持高度灵活的输入输出配置,Cosmos 3无缝融合视觉-语言模型、视频生成器、世界模拟器与世界-动作模型,形成单一框架。评估显示,Cosmos 3在多样化的理解与生成任务上达到新基准,证实多模态世界模型可作为可扩展、通用的具身智能基础。经后训练的Cosmos 3模型在撰写技术报告时,被Artificial Analysis评为最佳开源文本到图像模型,被RoboArena评为最佳策略模型。为加速物理智能领域的开放研究与部署,代码、模型检查点、精选合成数据集及评估基准已按Linux基金会OpenMDW-1.1许可证公开,详见https://github.com/nvidia/cosmos 和 https://huggingface.co/collections/nvidia/cosmos3。项目官网:https://research.nvidia.com/labs/cosmos-lab/cosmos3。

原文摘要 · Abstract (English)

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.

多模态物理智能世界模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。