arXiv:2501.10958cs.CV2025-01中稿 · ICASSP 2025被引 4

提出早融合框架,提升低光照图像分割效率

Rethinking Early-Fusion Strategies for Improved Multimodal Image Segmentation

  • 采用早融合策略与简单特征聚类,减少参数和计算量
  • 在多个数据集上超越现有方法,参数和计算量更低
  • 适合资源受限场景下的多模态图像分割应用

RGB与热成像融合在低光照条件下具有提升语义分割性能的潜力。现有方法通常采用双分支编码器进行多模态特征提取,并设计复杂的特征融合策略,但这类方法在特征提取与融合过程中需要大量参数更新和计算开销。为解决该问题,我们提出一种基于早融合策略的新型多模态融合网络(EFNet),结合简单有效的特征聚类机制,实现高效的RGB-T语义分割训练。此外,我们还设计了一种基于欧氏距离的轻量级多尺度特征聚合解码器。在多个数据集上的实验验证了方法的有效性,其性能优于先前最优方法,同时参数量和计算量均更低。

原文摘要 · Abstract (English)

RGB and thermal image fusion have great potential to exhibit improved semantic segmentation in low-illumination conditions. Existing methods typically employ a two-branch encoder framework for multimodal feature extraction and design complicated feature fusion strategies to achieve feature extraction and fusion for multimodal semantic segmentation. However, these methods require massive parameter updates and computational effort during the feature extraction and fusion. To address this issue, we propose a novel multimodal fusion network (EFNet) based on an early fusion strategy and a simple but effective feature clustering for training efficient RGB-T semantic segmentation. In addition, we also propose a lightweight and efficient multi-scale feature aggregation decoder based on Euclidean distance. We validate the effectiveness of our method on different datasets and outperform previous state-of-the-art methods with lower parameters and computation.

多模态图像分割早融合轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。