融合摄像头与激光雷达,实现跨区域、远距离交通标志的稳定识别
Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

- 用激光雷达与摄像头融合,利用几何不变性定位标志,摆脱视觉差异影响
- 在200米外小至10×10像素的标志仍能精准检测,误检率仅0.49%
- 新增属性分析模块,可判断遮挡程度、可读性与道路相关性,助力决策
可靠的交通标志检测是自动驾驶系统全球部署的前提,其性能需在不同国家、距离和天气下保持稳定。尽管已有进展,现有视觉方法仍存在三大缺陷:跨区域泛化能力差(各国标志差异大)、远距离小目标检测性能下降(200米处标志仅占10×10像素)、车辆接近时因强非线性透视畸变导致跟踪失效。本文提出多模态检测框架,结合相机与激光雷达(LiDAR)感知。引入强度感知可变形融合模块,将激光雷达的反光特征与图像特征对齐,基于几何不变量进行检测,而非依赖区域特定的视觉外观。进一步设计双运动模型追踪器,显式建模车辆接近过程中的非线性透视变化,显著提升时间一致性。此外,构建语义属性分类管道,估计遮挡程度、可读性、嵌入度及道路相关性,为下游规划提供可操作上下文。在覆盖60多个国家、超过2500小时驾驶数据的自建数据集上评估,该方案在221,068个序列中实现0.49%的物体漏检率(OMR),验证了商用级自动驾驶系统中全局泛化的交通标志感知能力。
原文摘要 · Abstract (English)
Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。