RoadMamba通过双分支融合局部纹理与全局语义,提升道路表面分类准确率。
RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification
- 采用双状态空间模型同时捕捉道路的局部纹理和全局语义特征。
- 在百万级样本数据集上达到当前最佳性能,显著优于现有Mamba架构。
- 适合自动驾驶中需要高精度路面识别的场景,尤其关注细节纹理的系统。
基于视觉技术提前获取道路表面状况,可为自动驾驶车辆的规划与控制系统提供有效信息,从而提升行车安全与舒适性。近期基于状态空间模型的Mamba架构在视觉任务中表现出色,得益于其高效的全局感受野。然而,现有Mamba架构因难以有效提取道路表面的局部纹理,在视觉道路表面分类任务中未能达到顶尖水平。本文首次探索了视觉Mamba架构在道路表面分类中的潜力,提出一种有效融合局部与全局感知的方法——RoadMamba。具体而言,利用双状态空间模型(DualSSM)分别提取道路表面的全局语义与局部纹理,并通过双注意力融合(DAF)进行特征解码与融合。此外,设计双辅助损失显式约束双分支,防止网络过度依赖深层大感受野的全局语义而忽略局部纹理。实验表明,所提出的RoadMamba在包含100万样本的大规模道路表面分类数据集上达到了当前最优性能。
原文摘要 · Abstract (English)
Acquiring the road surface conditions in advance based on visual technologies provides effective information for the planning and control system of autonomous vehicles, thus improving the safety and driving comfort of the vehicles. Recently, the Mamba architecture based on state-space models has shown remarkable performance in visual processing tasks, benefiting from the efficient global receptive field. However, existing Mamba architectures struggle to achieve state-of-the-art visual road surface classification due to their lack of effective extraction of the local texture of the road surface. In this paper, we explore for the first time the potential of visual Mamba architectures for road surface classification task and propose a method that effectively combines local and global perception, called RoadMamba. Specifically, we utilize the Dual State Space Model (DualSSM) to effectively extract the global semantics and local texture of the road surface and decode and fuse the dual features through the Dual Attention Fusion (DAF). In addition, we propose a dual auxiliary loss to explicitly constrain dual branches, preventing the network from relying only on global semantic information from the deep large receptive field and ignoring the local texture. The proposed RoadMamba achieves the state-of-the-art performance in experiments on a large-scale road surface classification dataset containing 1 million samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。