让AI像人一样理解物理世界并做出合理动作决策
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- 用分层与二维本体论建模物理常识和具身推理
- 7B和56B模型在物理常识与具身任务上显著提升
- 适合研究具身智能、物理推理与多模态大模型的开发者
物理人工智能系统需在真实世界中感知、理解并执行复杂动作。本文提出Cosmos-Reason1模型,通过长链式思维过程,以自然语言理解物理世界并生成恰当的具身决策(如下一步动作)。我们定义了物理常识与具身推理的核心能力:采用分层本体表示空间、时间与物理基础知识;使用二维本体实现跨不同物理形态的泛化。基于此,构建两个多模态大模型——Cosmos-Reason1-7B与Cosmos-Reason1-56B。训练分为两阶段:物理人工智能监督微调(SFT)与强化学习(RL)。为评估模型,我们依据本体构建综合性基准测试。结果表明,物理AI SFT与RL带来显著性能提升。为促进物理人工智能发展,代码与预训练模型已按NVIDIA开放模型许可证发布于https://github.com/nvidia-cosmos/cosmos-reason1。
原文摘要 · Abstract (English)
Physical AI systems need to perceive, understand, and perform complex actions in the physical world. In this paper, we present the Cosmos-Reason1 models that can understand the physical world and generate appropriate embodied decisions (e.g., next step action) in natural language through long chain-of-thought reasoning processes. We begin by defining key capabilities for Physical AI reasoning, with a focus on physical common sense and embodied reasoning. To represent physical common sense, we use a hierarchical ontology that captures fundamental knowledge about space, time, and physics. For embodied reasoning, we rely on a two-dimensional ontology that generalizes across different physical embodiments. Building on these capabilities, we develop two multimodal large language models, Cosmos-Reason1-7B and Cosmos-Reason1-56B. We curate data and train our models in two stages: Physical AI supervised fine-tuning (SFT) and Physical AI reinforcement learning (RL). To evaluate our models, we build comprehensive benchmarks for physical common sense and embodied reasoning according to our ontologies. Evaluation results show that Physical AI SFT and RL bring significant improvements. To facilitate the development of Physical AI, we make our code and pre-trained models available under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-reason1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。