用视觉Transformer融合摄像头与激光雷达数据,提升复杂路况下的交通物体分割精度
A Novel Vision Transformer for Camera-LiDAR Fusion based Traffic Object Segmentation
- 基于视觉Transformer的多模态融合架构,利用自注意力机制对图像和点云数据联合建模
- 在多种天气条件下实现对行人、骑行者、交通标志等物体的准确分割,表现优于现有方法
- 适合自动驾驶感知系统研发人员,尤其关注多传感器融合与复杂环境鲁棒性提升
本文提出相机-激光雷达融合视觉变压器(CLFT)模型,用于交通物体分割,通过视觉变压器融合摄像头与激光雷达数据。基于视觉变压器的自注意力机制,该模型扩展了对行人、骑行者、交通标志等多样化目标的分割能力,并在不同天气条件下表现出良好性能。尽管如此,模型在黑暗和雨天等恶劣环境下仍存在挑战,凸显进一步优化的必要性。总体而言,CLFT为自动驾驶感知提供了有力解决方案,推动了多模态融合与物体分割的技术前沿,但需持续改进以充分释放其在实际部署中的潜力。
原文摘要 · Abstract (English)
This paper presents Camera-LiDAR Fusion Transformer (CLFT) models for traffic object segmentation, which leverage the fusion of camera and LiDAR data using vision transformers. Building on the methodology of visual transformers that exploit the self-attention mechanism, we extend segmentation capabilities with additional classification options to a diverse class of objects including cyclists, traffic signs, and pedestrians across diverse weather conditions. Despite good performance, the models face challenges under adverse conditions which underscores the need for further optimization to enhance performance in darkness and rain. In summary, the CLFT models offer a compelling solution for autonomous driving perception, advancing the state-of-the-art in multimodal fusion and object segmentation, with ongoing efforts required to address existing limitations and fully harness their potential in practical deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。