对比12种深度学习模型,提升航拍图中车道线提取精度。
Advancements in Road Lane Mapping: Comparative Fine-Tuning Analysis of Deep Learning-based Semantic Segmentation Methods Using Aerial Imagery
- 用部分标注的水牛城数据集微调预训练模型。
- 微调后平均交并比达33.56%至76.11%,召回率66.0%~98.96%。
- 基于Transformer的模型优于传统卷积网络,适合自动驾驶地图构建。
本研究针对自动驾驶车辆所需高精地图中的车道信息提取问题,利用航拍影像开展深度学习语义分割模型的对比分析。尽管地球观测数据为地图生成提供丰富资源,但遥感领域针对车道线提取的专用模型仍较欠缺。本文系统比较了12种主流深度学习语义分割模型在高清遥感图像中的车道线提取性能,采用部分标注的Waterloo Urban Scene数据集进行微调,并以SkyScapes数据集预训练,模拟真实部署场景下的部分标注条件。实验表明,微调后模型性能显著提升,平均交并比(mean IoU)在33.56%至76.11%之间,召回率介于66.0%至98.96%。基于Transformer的模型整体表现优于卷积神经网络,凸显了预训练与微调对高精地图开发的关键作用。
原文摘要 · Abstract (English)
This research addresses the need for high-definition (HD) maps for autonomous vehicles (AVs), focusing on road lane information derived from aerial imagery. While Earth observation data offers valuable resources for map creation, specialized models for road lane extraction are still underdeveloped in remote sensing. In this study, we perform an extensive comparison of twelve foundational deep learning-based semantic segmentation models for road lane marking extraction from high-definition remote sensing images, assessing their performance under transfer learning with partially labeled datasets. These models were fine-tuned on the partially labeled Waterloo Urban Scene dataset, and pre-trained on the SkyScapes dataset, simulating a likely scenario of real-life model deployment under partial labeling. We observed and assessed the fine-tuning performance and overall performance. Models showed significant performance improvements after fine-tuning, with mean IoU scores ranging from 33.56% to 76.11%, and recall ranging from 66.0% to 98.96%. Transformer-based models outperformed convolutional neural networks, emphasizing the importance of model pre-training and fine-tuning in enhancing HD map development for AV navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。