用多步推理让AI跨国家精准识别路标,零样本也能行。
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
- 分三步引导大模型:定位、特征对比、差异辨别
- 跨国家路标识别准确率达85%~97%
- 无需训练数据,适合真实复杂道路场景
本研究提出跨域多步思维(CdMT)框架,以提升野外零样本细粒度交通标志识别性能。由于清洁模板标志与真实道路标志存在跨域差异,尤其在跨国场景下标志样式差异显著,现有方法表现受限。CdMT利用大模态模型的多步推理能力,设计上下文、特征和差异三类描述来构建多重推理流程。上下文描述通过中心坐标提示优化,实现复杂路况中目标标志的精准定位,并基于先验假设过滤无关响应;特征描述借助模板标志的上下文学习,弥合跨域差距,增强细粒度识别;差异描述则通过区分相似标志间的细微差别,强化多模态推理能力。该方法不依赖训练数据,仅需统一指令即可实现跨国识别。在三个基准数据集及两个不同国家的真实数据集上测试,CdMT在GTSRB、BTSD、TT-100K、Sapporo和Yokohama数据集上的识别准确率分别为0.93、0.89、0.97、0.89和0.85,优于所有现有方法。
原文摘要 · Abstract (English)
In this study, we propose Cross-domain Multi-step Thinking (CdMT) to improve zero-shot fine-grained traffic sign recognition (TSR) performance in the wild. Zero-shot fine-grained TSR in the wild is challenging due to the cross-domain problem between clean template traffic signs and real-world counterparts, and existing approaches particularly struggle with cross-country TSR scenarios, where traffic signs typically differ between countries. The proposed CdMT framework tackles these challenges by leveraging the multi-step reasoning capabilities of large multimodal models (LMMs). We introduce context, characteristic, and differential descriptions to design multiple thinking processes for LMMs. Context descriptions, which are enhanced by center coordinate prompt optimization, enable the precise localization of target traffic signs in complex road images and filter irrelevant responses via novel prior traffic sign hypotheses. Characteristic descriptions, which are derived from in-context learning with template traffic signs, bridge cross-domain gaps and enhance fine-grained TSR. Differential descriptions refine the multimodal reasoning ability of LMMs by distinguishing subtle differences among similar signs. CdMT is independent of training data and requires only simple and uniform instructions, enabling it to achieve cross-country TSR. We conducted extensive experiments on three benchmark datasets and two real-world datasets from different countries. The proposed CdMT framework achieved superior performance compared with other state-of-the-art methods on all five datasets, with recognition accuracies of 0.93, 0.89, 0.97, 0.89, and 0.85 on the GTSRB, BTSD, TT-100K, Sapporo, and Yokohama datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。