融合激光雷达与摄像头数据,无需视频训练即可提升全景分割精度
LiDAR-Camera Fusion for Video Panoptic Segmentation without Video Training
- 设计轻量级特征融合模块,实现图像与激光雷达数据高效结合
- 在图像和视频全景分割上均提升最高5个百分点
- 仅需两个简单修改,适合部署于自动驾驶系统
全景分割将实例分割与语义分割结合,因对场景的完整表征而在自动驾驶领域受到广泛关注。该任务可应用于摄像头和激光雷达传感器,但针对两者融合以增强图像全景分割的研究仍有限。尽管已有研究指出三维数据对基于摄像头的场景感知有益,却未专门探讨其对图像及视频全景分割(VPS)的影响。本文提出一种特征融合模块,通过融合激光雷达与图像数据来提升全景分割性能。此外,我们证明,在此基础上仅进行两项简单改进,即可在不使用视频训练数据的情况下,进一步显著提升视频全景分割质量。实验结果表明,图像与视频全景分割的评估指标均有显著提升,最高达5个百分点。
原文摘要 · Abstract (English)
Panoptic segmentation, which combines instance and semantic segmentation, has gained a lot of attention in autonomous vehicles, due to its comprehensive representation of the scene. This task can be applied for cameras and LiDAR sensors, but there has been a limited focus on combining both sensors to enhance image panoptic segmentation (PS). Although previous research has acknowledged the benefit of 3D data on camera-based scene perception, no specific study has explored the influence of 3D data on image and video panoptic segmentation (VPS).This work seeks to introduce a feature fusion module that enhances PS and VPS by fusing LiDAR and image data for autonomous vehicles. We also illustrate that, in addition to this fusion, our proposed model, which utilizes two simple modifications, can further deliver even more high-quality VPS without being trained on video data. The results demonstrate a substantial improvement in both the image and video panoptic segmentation evaluation metrics by up to 5 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。