提出3D驾驶场景分块技术,统一多视角重建与理解
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
- 用3D可变形交叉注意力将视觉特征转为场景令牌
- 在nuScenes上实现多任务性能优于现有方法
- 适合自动驾驶多模态模型的高效视觉接口
随着视觉-语言-动作模型和世界模型在自动驾驶中的应用日益广泛,可扩展的图像分块技术成为视觉模态的关键接口。然而,现有分块器多针对单目2D场景,应用于高分辨率多视角驾驶场景时效率低且视角间不一致。为此,我们提出DriveTok,一种高效的3D驾驶场景分块方法,用于统一多视角重建与理解。DriveTok首先从视觉基础模型获取语义丰富的视觉特征,再通过3D可变形交叉注意力将其转换为场景令牌。解码阶段,采用多视角Transformer从场景令牌重构多视角特征,并通过多个头实现RGB、深度与语义重建。同时,在场景令牌上直接添加3D头,用于3D语义占据预测,增强空间感知能力。通过多任务训练目标,DriveTok学习到融合语义、几何与纹理信息的统一场景令牌,实现高效多视角分块。在广泛使用的nuScenes数据集上的大量实验表明,DriveTok生成的场景令牌在图像重建、语义分割、深度预测及3D占据预测任务中均表现优异。
原文摘要 · Abstract (English)
With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing tokenizers are designed for monocular and 2D scenes, leading to inefficiency and inter-view inconsistency when applied to high-resolution multi-view driving scenes. To address this, we propose DriveTok, an efficient 3D driving scene tokenizer for unified multi-view reconstruction and understanding. DriveTok first obtains semantically rich visual features from vision foundation models and then transforms them into the scene tokens with 3D deformable cross-attention. For decoding, we employ a multi-view transformer to reconstruct multi-view features from the scene tokens and use multiple heads to obtain RGB, depth, and semantic reconstructions. We also add a 3D head directly on the scene tokens for 3D semantic occupancy prediction for better spatial awareness. With the multiple training objectives, DriveTok learns unified scene tokens that integrate semantic, geometric, and textural information for efficient multi-view tokenization. Extensive experiments on the widely used nuScenes dataset demonstrate that the scene tokens from DriveTok perform well on image reconstruction, semantic segmentation, depth prediction, and 3D occupancy prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。