arXiv:2409.15077cs.CV2024-09ICRA被引 6

用CLIP模型提升全球交通标志识别鲁棒性

TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition

  • 基于提示工程生成针对交通标志的文本描述
  • 通过动态加权融合实现跨区域性能提升
  • 首个用于全球交通标志识别的CLIP方法

交通标志是导航与交通管控的关键地图要素。然而,现有交通标志识别方法依赖传统深度学习模型,在不同区域间数据分布差异下性能显著下降。本文提出TSCLIP,一种基于对比语言图像预训练(CLIP)模型的鲁棒微调方法,用于全球跨区域交通标志识别。我们首先整合来自十个不同来源的数据,构建跨区域交通标志基准数据集。随后设计针对交通标志特性的提示工程方案,结合具体场景描述和对应规则生成目标文本描述。在TSCLIP微调过程中,采用自适应动态权重集成(ADWE)机制,将每轮训练结果与零样本CLIP模型输出无缝融合,确保模型在获取新知识的同时保持泛化能力。据作者所知,TSCLIP是首个应用于全球跨区域交通标志识别任务的对比语言图像模型。项目主页见:https://github.com/guoyangzhao/TSCLIP。

原文摘要 · Abstract (English)

Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, which typically suffer from significant performance degradation considering the variations in data distribution across different regions. In this paper, we propose TSCLIP, a robust fine-tuning approach with the contrastive language-image pre-training (CLIP) model for worldwide cross-regional traffic sign recognition. We first curate a cross-regional traffic sign benchmark dataset by combining data from ten different sources. Then, we propose a prompt engineering scheme tailored to the characteristics of traffic signs, which involves specific scene descriptions and corresponding rules to generate targeted text descriptions. During the TSCLIP fine-tuning process, we implement adaptive dynamic weight ensembling (ADWE) to seamlessly incorporate outcomes from each training iteration with the zero-shot CLIP model. This approach ensures that the model retains its ability to generalize while acquiring new knowledge about traffic signs. To the best knowledge of authors, TSCLIP is the first contrastive language-image model used for the worldwide cross-regional traffic sign recognition task. The project website is available at: https://github.com/guoyangzhao/TSCLIP.

交通标志CLIP跨区域识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。