通过融合实例表示提升LSS-BEV感知的几何建模能力
LSSInst: Improving Geometric Modeling in LSS-Based BEV Perception with Instance Representation
- 提出两阶段检测框架,联合使用BEV与实例表示
- 在nuScenes上显著提升现有LSS方法性能,无需额外组件
- 设计实例适配器解决空间表征差异,增强几何细节保留
随着纯摄像头3D目标检测在自动驾驶中的兴起,基于鸟瞰图(BEV)表示的方法,特别是源自前视图变换范式的升维-投射-射击(LSS)框架,近年来取得显著进展。基于深度分布预测构建的视锥体BEV表示,适合从多视角图像中学习道路结构和场景布局。然而,为保持计算效率,压缩后的BEV表示(如分辨率和轴向维度)不可避免地削弱了个体几何细节的保留,影响方法的通用性和适用性。为此,我们提出LSSInst,一种结合BEV与实例表示的两阶段目标检测器。该检测器利用细粒度像素级特征,可灵活集成到现有LSS-based BEV网络中。由于两种表示空间存在固有差异,我们设计实例适配器以实现BEV到实例的语义一致性,而非简单传递提议。大量实验表明,所提框架具有优异的泛化能力和性能,在无需额外组件的情况下提升现代LSS-based BEV感知方法的表现,并在大规模nuScenes基准上超越当前SOTA方法。
原文摘要 · Abstract (English)
With the attention gained by camera-only 3D object detection in autonomous driving, methods based on Bird-Eye-View (BEV) representation especially derived from the forward view transformation paradigm, i.e., lift-splat-shoot (LSS), have recently seen significant progress. The BEV representation formulated by the frustum based on depth distribution prediction is ideal for learning the road structure and scene layout from multi-view images. However, to retain computational efficiency, the compressed BEV representation such as in resolution and axis is inevitably weak in retaining the individual geometric details, undermining the methodological generality and applicability. With this in mind, to compensate for the missing details and utilize multi-view geometry constraints, we propose LSSInst, a two-stage object detector incorporating BEV and instance representations in tandem. The proposed detector exploits fine-grained pixel-level features that can be flexibly integrated into existing LSS-based BEV networks. Having said that, due to the inherent gap between two representation spaces, we design the instance adaptor for the BEV-to-instance semantic coherence rather than pass the proposal naively. Extensive experiments demonstrated that our proposed framework is of excellent generalization ability and performance, which boosts the performances of modern LSS-based BEV perception methods without bells and whistles and outperforms current LSS-based state-of-the-art works on the large-scale nuScenes benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。