用查询生成3D高斯分布,提升自动驾驶稀疏感知模型性能
SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
- 通过自监督拼贴重建多视角图像和深度图,学习精细上下文特征
- 在占用预测上提升1.3 mIoU,3D检测上提升1.0 NDS
- 适用于需要高效推理的自动驾驶感知任务
稀疏感知模型(SPMs)采用查询驱动范式,跳过显式的鸟瞰图或体素构建,实现高效计算与快速推理。本文提出SQS,一种专为自动驾驶设计的查询式拼贴预训练方法。SQS引入可插拔模块,在预训练阶段从稀疏查询中预测3D高斯表示,利用自监督拼贴技术通过多视角图像与深度图重建学习细粒度上下文特征。微调时,预训练的高斯查询通过查询交互机制无缝融入下游网络,明确连接预训练查询与任务特定查询,有效适配占用预测与3D目标检测的多样化需求。在多个自动驾驶基准上的实验表明,SQS在多种基于查询的3D感知任务中均取得显著性能提升,尤其在占用预测与3D目标检测上大幅超越现有最优预训练方法(如占用预测提升1.3 mIoU,3D检测提升1.0 NDS)。
原文摘要 · Abstract (English)
Sparse Perception Models (SPMs) adopt a query-driven paradigm that forgoes explicit dense BEV or volumetric construction, enabling highly efficient computation and accelerated inference. In this paper, we introduce SQS, a novel query-based splatting pre-training specifically designed to advance SPMs in autonomous driving. SQS introduces a plug-in module that predicts 3D Gaussian representations from sparse queries during pre-training, leveraging self-supervised splatting to learn fine-grained contextual features through the reconstruction of multi-view images and depth maps. During fine-tuning, the pre-trained Gaussian queries are seamlessly integrated into downstream networks via query interaction mechanisms that explicitly connect pre-trained queries with task-specific queries, effectively accommodating the diverse requirements of occupancy prediction and 3D object detection. Extensive experiments on autonomous driving benchmarks demonstrate that SQS delivers considerable performance gains across multiple query-based 3D perception tasks, notably in occupancy prediction and 3D object detection, outperforming prior state-of-the-art pre-training approaches by a significant margin (i.e., +1.3 mIoU on occupancy prediction and +1.0 NDS on 3D detection).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。