构建百万级交通标志数据集,评估模型在真实场景下的鲁棒性表现。
Traffic Sign Recognition in Autonomous Driving: Dataset, Benchmark, and Field Experiment
- 提出跨区域、罕见类别等五类挑战场景,诊断模型性能边界。
- 发现语义对齐显著提升跨区域识别与稀有类别检测能力。
- 适合自动驾驶感知系统研发者及模型鲁棒性研究者参考。
交通标志识别(TSR)是自动驾驶核心感知能力,需应对跨区域差异、长尾类别和语义模糊等实际挑战。现有数据集与基准测试难以揭示不同模型在这些条件下的表现。本文发布TS-1M,一个包含超过一百万张真实世界图像、覆盖454个标准化类别的大规模全球多样性数据集,并设计诊断性基准评估模型能力边界。除常规训练-测试外,还提供跨区域识别、罕见类别识别、低清晰度鲁棒性、语义文本理解等挑战设置,实现对现代TSR模型的系统性细粒度评估。基于TS-1M,我们对比了三类代表性学习范式:经典监督模型、自监督预训练模型和多模态视觉-语言模型(VLMs)。分析表明,语义对齐是跨区域泛化与稀有类别识别的关键,而纯视觉模型仍易受外观变化与数据不平衡影响。最后,通过真实道路自动驾驶实验验证了TS-1M的实际价值,将交通标志识别与语义推理、空间定位结合,支持地图级决策约束。总体而言,TS-1M建立了一个参考级诊断基准,为鲁棒且语义感知的交通标志感知提供了原则性洞见。
原文摘要 · Abstract (English)
Traffic Sign Recognition (TSR) is a core perception capability for autonomous driving, where robustness to cross-region variation, long-tailed categories, and semantic ambiguity is essential for reliable real-world deployment. Despite steady progress in recognition accuracy, existing traffic sign datasets and benchmarks offer limited diagnostic insight into how different modeling paradigms behave under these practical challenges. We present TS-1M, a large-scale and globally diverse traffic sign dataset comprising over one million real-world images across 454 standardized categories, together with a diagnostic benchmark designed to analyze model capability boundaries. Beyond standard train-test evaluation, we provide a suite of challenge-oriented settings, including cross-region recognition, rare-class identification, low-clarity robustness, and semantic text understanding, enabling systematic and fine-grained assessment of modern TSR models. Using TS-1M, we conduct a unified benchmark across three representative learning paradigms: classical supervised models, self-supervised pretrained models, and multimodal vision-language models (VLMs). Our analysis reveals consistent paradigm-dependent behaviors, showing that semantic alignment is a key factor for cross-region generalization and rare-category recognition, while purely visual models remain sensitive to appearance shift and data imbalance. Finally, we validate the practical relevance of TS-1M through real-scene autonomous driving experiments, where traffic sign recognition is integrated with semantic reasoning and spatial localization to support map-level decision constraints. Overall, TS-1M establishes a reference-level diagnostic benchmark for TSR and provides principled insights into robust and semantic-aware traffic sign perception. Project page: https://guoyangzhao.github.io/projects/ts1m.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。