arXiv:2503.14492cs.CVcs.AI2025-03被引 92

用多种视觉输入精准控制虚拟世界生成,适合机器人与自动驾驶仿真

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

  • 根据分割、深度、边缘等多模态输入自适应调节控制强度
  • 实现高精度世界模拟,支持物理人工智能中的真实世界迁移
  • 可实现实时生成,适合机器人和自动驾驶数据增强

我们提出 Cosmos-Transfer,一种基于多空间模态输入(如分割图、深度图、边缘图)的条件世界生成模型。该模型采用自适应的空间条件机制,可在不同空间位置灵活调整各类输入的权重,实现高度可控的世界生成。该方法适用于多种世界到世界的迁移场景,包括 Sim2Real。通过大量评估验证了其在物理人工智能中的应用,涵盖机器人 Sim2Real 和自动驾驶数据增强。此外,我们还提出一种推理扩展策略,可在 NVIDIA GB200 NVL72 机架上实现实时世界生成。为推动领域发展,我们已开源模型与代码至 https://github.com/nvidia-cosmos/cosmos-transfer1。

原文摘要 · Abstract (English)

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.

世界生成多模态控制物理AISim2Real

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。