arXiv:2507.22477cs.CVcs.AI2025-07中稿 · ACM MM 2025被引 4

轻量级多模态裂缝分割模型,高效融合形态与纹理信息。

LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

  • 通过自适应感知模块动态捕捉多模态裂缝特征
  • 仅5.35M参数即达F1 0.8204、mIoU 0.8465
  • 适合资源受限场景下的高精度裂缝检测

利用多模态数据实现低计算成本的像素级裂缝分割仍是关键挑战。现有方法缺乏跨模态特征的自适应感知与高效交互融合能力。为此,我们提出轻量级自适应提示感知视觉状态空间网络(LIDAR),在多模态裂缝场景下高效感知并融合形态与纹理线索,生成清晰的像素级裂缝分割图。LIDAR由轻量级自适应提示感知视觉状态空间模块(LacaVSS)和轻量级双域动态协同融合模块(LD3CF)组成。LacaVSS通过掩码引导的高效动态扫描策略(EDG-SS)自适应建模裂缝提示;LD3CF结合自适应频域感知器(AFDP)与双池化融合策略,有效捕捉跨模态的空间与频域特征。此外,设计轻量级动态调制多核卷积(LDMK),以极低开销感知复杂形态结构,替代大部分卷积操作。在三个数据集上的实验表明,该方法优于其他最新方法。在光场深度数据集上,仅用5.35M参数即达到F1 0.8204、mIoU 0.8465。代码与数据集见https://github.com/Karl1109/LIDAR-Mamba。

原文摘要 · Abstract (English)

Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba.

裂缝分割多模态融合轻量模型视觉状态空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。