提出CoIn3D框架,让多摄像头3D检测模型跨配置通用
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
- 通过空间感知特征调制和相机感知数据增强,显式建模不同相机配置的几何差异
- 在NuScenes、Waymo、Lyft上跨配置检测性能显著提升,三种主流模型均适用
- 无需额外训练即可生成新视角图像,适合部署于新型多摄像头设备的场景
多摄像头3D目标检测(MC3D)因机器人与自动驾驶等多传感器系统普及而受到关注。然而,现有模型难以泛化到未见的多摄像头配置。当前方法仅用统一元相机表征,缺乏全面考虑。本文重新审视该问题,发现源配置与目标配置间存在空间先验差异,包括内参、外参及阵列布局的不同。为此,提出CoIn3D框架,实现从源配置到未见目标配置的强泛化能力。CoIn3D通过空间感知特征调制(SFM)和相机感知数据增强(CDA),分别在特征嵌入和图像观测中显式融合所有空间先验。SFM整合焦距、地面深度、地面梯度与Plücker坐标四类空间表示;CDA采用免训练的动态新视角图像合成方案,提升不同配置下的观测多样性。大量实验表明,CoIn3D在NuScenes、Waymo、Lyft等基准数据集上,对代表BEVDepth、BEVFormer和PETR的三种主流MC3D范式均实现优异的跨配置性能。
原文摘要 · Abstract (English)
Multi-camera 3D object detection (MC3D) has attracted increasing attention with the growing deployment of multi-sensor physical agents, such as robots and autonomous vehicles. However, MC3D models still struggle to generalize to unseen platforms with new multi-camera configurations. Current solutions simply employ a meta-camera for unified representation but lack comprehensive consideration. In this paper, we revisit this issue and identify that the devil lies in spatial prior discrepancies across source and target configurations, including different intrinsics, extrinsics, and array layouts. To address this, we propose CoIn3D, a generalizable MC3D framework that enables strong transferability from source configurations to unseen target ones. CoIn3D explicitly incorporates all identified spatial priors into both feature embedding and image observation through spatial-aware feature modulation (SFM) and camera-aware data augmentation (CDA), respectively. SFM enriches feature space by integrating four spatial representations, such as focal length, ground depth, ground gradient, and Plücker coordinate. CDA improves observation diversity under various configurations via a training-free dynamic novel-view image synthesis scheme. Extensive experiments demonstrate that CoIn3D achieves strong cross-configuration performance on landmark datasets such as NuScenes, Waymo, and Lyft, under three dominant MC3D paradigms represented by BEVDepth, BEVFormer, and PETR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。