基于脑启发机制的多轮交互网络,提升红外与可见光图像融合的路面语义分割精度。
BIMII-Net: Brain-Inspired Multi-Iterative Interactive Network for RGB-T Road Scene Semantic Segmentation
- 采用类脑连续耦合网络提取纹理与局部信息
- 设计跨模态显式注意力模块实现多层级特征融合
- 多层迭代解码结构协同优化细节与全局结构,适合自动驾驶场景
RGB-T道路场景语义分割通过融合可见光与热成像信息,提升复杂环境(如光照不足或遮挡)下的视觉理解能力。现有模型多依赖简单加法或拼接策略,忽略不同层次信息的差异。为此,本文提出脑启发多轮交互网络(BIMII-Net)。首先,基于类脑模型设计深度连续耦合神经网络(DCCNN),满足自动驾驶中对纹理与局部信息的精准提取需求。其次,在特征融合阶段引入跨模态显式注意力增强融合模块(CEAEF-Module),有效整合多层级特征。最后,构建互补式多层解码结构,包含浅层特征迭代模块(SFI-Module)、深层特征迭代模块(DFI-Module)与多特征增强模块(MFE-Module),协同提取纹理细节与全局骨架信息,并通过多模块联合监督优化分割结果。实验表明,BIMII-Net在脑启发计算领域达到当前最优性能,显著优于多数现有RGB-T语义分割方法,并在多个公开数据集上展现强泛化能力,验证了脑启发模型在多模态图像分割中的有效性。
原文摘要 · Abstract (English)
RGB-T road scene semantic segmentation enhances visual scene understanding in complex environments characterized by inadequate illumination or occlusion by fusing information from RGB and thermal images. Nevertheless, existing RGB-T semantic segmentation models typically depend on simple addition or concatenation strategies or ignore the differences between information at different levels. To address these issues, we proposed a novel RGB-T road scene semantic segmentation network called Brain-Inspired Multi-Iteration Interaction Network (BIMII-Net). First, to meet the requirements of accurate texture and local information extraction in road scenarios like autonomous driving, we proposed a deep continuous-coupled neural network (DCCNN) architecture based on a brain-inspired model. Second, to enhance the interaction and expression capabilities among multi-modal information, we designed a cross explicit attention-enhanced fusion module (CEAEF-Module) in the feature fusion stage of BIMII-Net to effectively integrate features at different levels. Finally, we constructed a complementary interactive multi-layer decoder structure, incorporating the shallow-level feature iteration module (SFI-Module), the deep-level feature iteration module (DFI-Module), and the multi-feature enhancement module (MFE-Module) to collaboratively extract texture details and global skeleton information, with multi-module joint supervision further optimizing the segmentation results. Experimental results demonstrate that BIMII-Net achieves state-of-the-art (SOTA) performance in the brain-inspired computing domain and outperforms most existing RGB-T semantic segmentation methods. It also exhibits strong generalization capabilities on multiple RGB-T datasets, proving the effectiveness of brain-inspired computer models in multi-modal image segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。