用世界模型生成高保真驾驶数据,解决罕见场景难获取问题
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
- 基于世界模型的可控生成框架,支持多视角、时序一致视频输出
- 生成数据显著缓解长尾分布问题,提升3D检测与驾驶策略泛化能力
- 开源工具链,适合自动驾驶感知与决策系统训练场景
为自动驾驶等安全关键物理人工智能系统收集真实世界数据耗时且成本高昂,尤其难以捕捉罕见边缘场景。为此,我们提出Cosmos-Drive-Dreams——一种合成数据生成(SDG)流水线,旨在生成具有挑战性的驾驶场景以支持感知与驾驶策略训练。该流水线由Cosmos-Drive驱动,其是一套基于NVIDIA Cosmos世界基础模型的专用驾驶领域模型,可实现可控、高保真、多视角、时空一致的驾驶视频生成。我们通过该方法扩展了驾驶数据集的数量与多样性,生成高质量、具挑战性的场景。实验表明,所生成数据有助于缓解长尾分布问题,并提升下游任务如3D车道检测、3D目标检测与驾驶策略学习的泛化性能。项目代码、数据集与模型权重已通过NVIDIA Cosmos平台开源。
原文摘要 · Abstract (English)
Collecting and annotating real-world data for safety-critical physical AI systems, such as Autonomous Vehicle (AV), is time-consuming and costly. It is especially challenging to capture rare edge cases, which play a critical role in training and testing of an AV system. To address this challenge, we introduce the Cosmos-Drive-Dreams - a synthetic data generation (SDG) pipeline that aims to generate challenging scenarios to facilitate downstream tasks such as perception and driving policy training. Powering this pipeline is Cosmos-Drive, a suite of models specialized from NVIDIA Cosmos world foundation model for the driving domain and are capable of controllable, high-fidelity, multi-view, and spatiotemporally consistent driving video generation. We showcase the utility of these models by applying Cosmos-Drive-Dreams to scale the quantity and diversity of driving datasets with high-fidelity and challenging scenarios. Experimentally, we demonstrate that our generated data helps in mitigating long-tail distribution problems and enhances generalization in downstream tasks such as 3D lane detection, 3D object detection and driving policy learning. We open source our pipeline toolkit, dataset and model weights through the NVIDIA's Cosmos platform. Project page: https://research.nvidia.com/labs/toronto-ai/cosmos_drive_dreams
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。