提出高效视图变换方法,提升多模态3D检测精度与速度
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
- 基于激光雷达引导生成自适应采样点和核函数,精准转换图像特征到鸟瞰图
- 在nuScenes数据集上达75.3% NDS,实现实时推理
- 适合追求高精度与低延迟的自动驾驶感知系统
鸟瞰图(BEV)表示下的多模态传感器融合已成为3D目标检测的主流方法。然而,现有方法常依赖深度估计算器或变压器编码器将图像特征转换至BEV空间,降低鲁棒性或引入显著计算开销。此外,视图变换中几何引导不足导致射线方向错位,限制了BEV表示的有效性。为此,本文提出高效视图变换(EVT),构建结构化更强的BEV表示,同时提升准确率与效率。首先,自适应采样与自适应投影(ASAP)利用激光雷达引导生成3D采样点和自适应核函数,更有效地将图像特征映射至BEV空间,并优化BEV表示。其次,改进的基于查询的检测框架结合分组混合查询选择与几何感知交叉注意力,有效捕捉物体共性及几何结构。在nuScenes测试集上,EVT达到75.3% NDS的领先性能,且支持实时推理。
原文摘要 · Abstract (English)
Multi-modal sensor fusion in Bird's Eye View (BEV) representation has become the leading approach for 3D object detection. However, existing methods often rely on depth estimators or transformer encoders to transform image features into BEV space, which reduces robustness or introduces significant computational overhead. Moreover, the insufficient geometric guidance in view transformation results in ray-directional misalignments, limiting the effectiveness of BEV representations. To address these challenges, we propose Efficient View Transformation (EVT), a novel 3D object detection framework that constructs a well-structured BEV representation, improving both accuracy and efficiency. Our approach focuses on two key aspects. First, Adaptive Sampling and Adaptive Projection (ASAP), which utilizes LiDAR guidance to generate 3D sampling points and adaptive kernels, enables more effective transformation of image features into BEV space and a refined BEV representation. Second, an improved query-based detection framework, incorporating group-wise mixed query selection and geometry-aware cross-attention, effectively captures both the common properties and the geometric structure of objects in the transformer decoder. On the nuScenes test set, EVT achieves state-of-the-art performance of 75.3% NDS with real-time inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。