用双潜空间模型实现自动驾驶中3D高斯表征的端到端预训练
DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
- 设计双潜空间世界模型,分两阶段实现高斯表征的自监督预训练
- 在nuScenes和SurroundOcc上,3D占用感知与运动规划性能显著提升
- 适合研究自动驾驶场景建模与多任务联合学习的开发者
基于视觉的自动驾驶因成本低、表现优而受到关注。相比密集鸟瞰图或稀疏查询模型,以3D语义高斯为核心的高斯中心方法是一种全面且稀疏的场景表征方式。本文提出DLWM,一种新型双潜空间世界模型范式,通过两阶段训练实现自动驾驶中高斯中心表征的端到端预训练。第一阶段,模型通过自监督重建多视角语义与深度图像,从查询中预测3D高斯;第二阶段,分别训练两个潜空间世界模型:一个用于下游占用感知与预测任务的高斯流引导潜变量预测,另一个用于运动规划的自身规划引导潜变量预测。在SurroundOcc与nuScenes基准上的大量实验表明,DLWM在3D占用感知、4D占用预测及运动规划任务中均取得显著性能提升。
原文摘要 · Abstract (English)
Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet sparse representation by describing scene with 3D semantic Gaussians. In this paper, we introduce DLWM, a novel paradigm with Dual Latent World Models specifically designed to enable holistic gaussian-centric pre-training in autonomous driving using two stages. In the first stage, DLWM predicts 3D Gaussians from queries by self-supervised reconstructing multi-view semantic and depth images. Equipped with fine-grained contextual features, in the second stage, two latent world models are trained separately for temporal feature learning, including Gaussian-flow-guided latent prediction for downstream occupancy perception and forecasting tasks, and ego-planning-guided latent prediction for motion planning. Extensive experiments in SurroundOcc and nuScenes benchmarks demonstrate that DLWM shows significant performance gains across Gaussian-centric 3D occupancy perception, 4D occupancy forecasting and motion planning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。