用动态专家选择提升复杂路况下交通标志识别精度与效率
Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios

- 构建分层解耦的专家池与轻量门控网络,实现图像级动态路由
- mAP50-95达76.8%,较基线提升2.3%,计算开销降低39.4%
- 适合需要高鲁棒性交通标志识别的自动驾驶系统部署
交通标志检测是自动驾驶与智能交通系统环境感知的基础。然而,现有检测器多依赖全局共享参数的静态推理,难以适应多样且非结构化的交通场景。单一静态模型常无法同时处理近距清晰样本与远距离小目标或恶劣天气等挑战性情况。为此,我们提出CBDES MoE TSR,一种分层解耦的异构混合专家(MoE)框架用于交通标志识别。该框架突破传统全局共享参数范式,引入异构的YOLO专家池与轻量门控网络,实现图像级动态路由。基于输入图像的语义特征,门控模块从专家池中选择最适配的专家模型,实现从固定参数拟合到按需动态表征的转变。该设计在提升特定场景特征提取能力的同时,保持可控的推理开销。实验表明,所提方法在复合交通标志数据集上实现了检测精度与效率的显著平衡:mAP50-95达76.8%,较基线(74.5%)提升2.3%,同时计算开销降低约39.4%。结果充分验证了该方法的有效性。
原文摘要 · Abstract (English)
Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, most existing detectors rely on static inference with globally shared parameters, limiting their ability to adapt to diverse and unstructured traffic scenarios. As a result, a single static model often struggles to simultaneously handle both clear near-range samples and challenging conditions such as distant small targets or adverse weather environments. To address this limitation, we propose CBDES MoE TSR, a hierarchically decoupled heterogeneous mixture-of-experts(MoE) framework for traffic sign recognition. The proposed framework departs from the conventional globally shared parameter paradigm by introducing a heterogeneous You Only Look Once (YOLO) expert pool together with a lightweight gating network, enabling an image-level dynamic routing mechanism. Based on the semantic characteristics of the input image, the gating module selectively activates the most suitable expert model from the expert pool, enabling a shift from fixed parameter fitting to on-demand dynamic representation. This design enhances feature extraction capability for specific scenarios while maintaining controlled inference overhead. Experimental results demonstrate that the proposed method achieves a remarkable balance between detection accuracy and efficiency on the composite traffic sign dataset. Specifically, our method attains an mAP50-95 of 76.8%, yielding a 2.3% improvement over the baseline method (74.5%) while simultaneously reducing computational overhead by approximately 39.4%. These findings robustly validate the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。