arXiv:2511.00062cs.CVcs.AI2025-11被引 171

用视频基础模型统一生成文本、图像和视频世界,提升物理智能仿真能力。

World Simulation with Video Foundation Models for Physical AI

论文配图:World Simulation with Video Foundation Models for Physical AI
图 1 · 摘自论文原文
  • 基于流模型架构,统一文本/图像/视频到世界的生成任务。
  • 在2亿条精选视频上训练,14B模型比前代视频质量与指令对齐显著提升。
  • 适合机器人、自动驾驶等需要高保真仿真与真实世界迁移的场景。

我们推出[Cosmos-Predict2.5],新一代用于物理智能的宇宙世界基础模型。基于流式架构,该模型统一实现文本到世界、图像到世界和视频到世界生成,并引入[Cosmos-Reason1]物理智能视觉语言模型,增强文本定位精度与世界模拟控制能力。模型在2亿条筛选视频上训练,并通过强化学习后训练优化,相比[.Cosmos-Predict1]在视频质量和指令对齐上均有显著提升,发布版本包含2B和14B参数规模。这些能力支持更可靠的合成数据生成、策略评估与闭环仿真,适用于机器人与自动驾驶系统。我们进一步推出[.Cosmos-Transfer2.5],一种类控制网框架,用于Sim2Real与Real2Real世界转换。尽管其规模仅为[.Cosmos-Transfer1]的1/3.5,却实现更高保真度与更长时序视频生成。二者共同构成可扩展具身智能的通用工具。为加速物理智能研究与部署,我们以NVIDIA开放模型许可证开源代码、预训练权重与精选基准,地址见GitHub。

原文摘要 · Abstract (English)

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single model and leverages [Cosmos-Reason1], a Physical AI vision-language model, to provide richer text grounding and finer control of world simulation. Trained on 200M curated video clips and refined with reinforcement learning-based post-training, [Cosmos-Predict2.5] achieves substantial improvements over [Cosmos-Predict1] in video quality and instruction alignment, with models released at 2B and 14B scales. These capabilities enable more reliable synthetic data generation, policy evaluation, and closed-loop simulation for robotics and autonomous systems. We further extend the family with [Cosmos-Transfer2.5], a control-net style framework for Sim2Real and Real2Real world translation. Despite being 3.5$\times$ smaller than [Cosmos-Transfer1], it delivers higher fidelity and robust long-horizon video generation. Together, these advances establish [Cosmos-Predict2.5] and [Cosmos-Transfer2.5] as versatile tools for scaling embodied intelligence. To accelerate research and deployment in Physical AI, we release source code, pretrained checkpoints, and curated benchmarks under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-predict2.5 and https://github.com/nvidia-cosmos/cosmos-transfer2.5. We hope these open resources lower the barrier to adoption and foster innovation in building the next generation of embodied intelligence.

世界模型物理智能视频生成仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。