用激光雷达指导单全景相机实现高效鸟瞰图分割,降低成本。
KD360-VoxelBEV: LiDAR and 360-degree Camera Cross Modality Knowledge Distillation for Bird's-Eye-View Segmentation
- 融合距离、强度和环境光的新型激光雷达表示,配合体素对齐视觉变换器。
- 学生模型在Dur360BEV上达25.6% IoU提升,推理速度31.2 FPS。
- 适用于多种摄像头布局,适合低成本自动驾驶部署。
我们提出首个专为单全景相机鸟瞰图(BEV)分割设计的跨模态知识蒸馏框架。该方法利用从距离、强度和环境通道融合生成的新型激光雷达图像表示,并结合体素对齐的视图变换器,在保持空间保真度的同时实现高效的BEV处理。训练时,一个高容量的激光雷达与相机融合教师网络提取丰富的空间与语义特征,用于向仅依赖单个360度全景相机图像的轻量级学生网络进行跨模态知识蒸馏。在Dur360BEV数据集上的大量实验表明,教师模型显著优于现有基于相机的BEV分割方法,实现25.6%的IoU提升。同时,蒸馏后的学生网络达到具有竞争力的性能,获得8.5%的IoU增益,并实现31.2 FPS的领先推理速度。此外,在KITTI-360(两个鱼眼相机)上的评估验证了该蒸馏框架可泛化至不同相机配置,凸显其可行性和鲁棒性。该方法降低了传感器复杂度与部署成本,为真实世界自动驾驶中的高效、低成本BEV分割提供了实用方案。
原文摘要 · Abstract (English)
We present the first cross-modality distillation framework specifically tailored for single-panoramic-camera Bird's-Eye-View (BEV) segmentation. Our approach leverages a novel LiDAR image representation fused from range, intensity and ambient channels, together with a voxel-aligned view transformer that preserves spatial fidelity while enabling efficient BEV processing. During training, a high-capacity LiDAR and camera fusion Teacher network extracts both rich spatial and semantic features for cross-modality knowledge distillation into a lightweight Student network that relies solely on a single 360-degree panoramic camera image. Extensive experiments on the Dur360BEV dataset demonstrate that our teacher model significantly outperforms existing camera-based BEV segmentation methods, achieving a 25.6\% IoU improvement. Meanwhile, the distilled Student network attains competitive performance with an 8.5\% IoU gain and state-of-the-art inference speed of 31.2 FPS. Moreover, evaluations on KITTI-360 (two fisheye cameras) confirm that our distillation framework generalises to diverse camera setups, underscoring its feasibility and robustness. This approach reduces sensor complexity and deployment costs while providing a practical solution for efficient, low-cost BEV segmentation in real-world autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。