arXiv:2606.08906cs.CV2026-06被引 4

提出DifferSeg框架,提升多模态图像分割的多样性与精度

DifferSeg: Towards Diverse Multimodal Binary Segmentation via Differential Perception and Frequency Guidance

论文配图:DifferSeg: Towards Diverse Multimodal Binary Segmentation via Differential Perception and Frequency Guidance
图 1 · 摘自论文原文
  • 通过可学习差分算子自适应对齐多模态特征,增强互补性
  • 设计频率引导解码器,兼顾高频细节与低频语义,边界更清晰
  • 适用于自然与医学图像,29个数据集上超越67种先进方法

在众多二值分割任务中,现有多模态方法通常采用固定特征拼接进行跨模态交互,并依赖以低频语义为主的简单解码器。然而,它们忽略了两个关键挑战:一是缺乏应对模态差异与互补性的自适应机制;二是缺少高效解码策略来平衡高低频表征。为此,本文提出一个简洁且通用的多模态二值分割框架DifferSeg,以同时解决上述问题。DifferSeg引入差分感知融合(DPF)模块,利用可学习差分算子自适应对齐多模态特征,通过残差融合增强互补性,有效缓解模态不匹配与融合冗余。此外,设计频率引导解码器(FGD),构建跨频率交互与多路径上采样,确保细节高频结构与语义低频表示的一致性,实现精细边界恢复与噪声抑制。得益于这些设计,DifferSeg可轻松泛化至多种二值分割任务,涵盖自然与医学模态。无需复杂组件,在29个公开数据集、18项下游任务中持续超越67种先进方法,展现卓越泛化能力与分割精度。代码与预训练模型将公开。

原文摘要 · Abstract (English)

In many binary segmentation tasks, most multimodal methods rely on fixed feature concatenation for cross-modal interaction and straightforward decoder designs dominated by low-frequency semantics. %ToDO: % However, they ignore two key challenges: one is the lack of an adaptive mechanism to handle modality discrepancies and complementarity, and the other is the absence of an efficient decoding strategy to balance both high- and low-frequency representations. % In this work, we propose a simple yet general multimodal binary segmentation framework, termed DifferSeg, to address both problems simultaneously. With the help of the differential perception fusion (DPF) module, DifferSeg employs learnable differential operators to adaptively align multimodal features and enhance their complementarity through residual fusion, effectively mitigating modality mismatch and fusion redundancy. % In addition, we design a frequency-guided decoder (FGD) that builds cross-frequency interactions and multi-path upsampling to maintain consistency between detailed high-frequency structures and semantic low-frequency representations, ensuring fine-grained boundary recovery and noise suppression. % Benefiting from these designs, DifferSeg can be easily generalized to diverse binary segmentation tasks, including both natural and medical modalities. Without bells and whistles, it consistently surpasses 67 state-of-the-art methods across 29 public datasets involving 18 downstream tasks, demonstrating superior generalization and segmentation accuracy.Code and pretrained models will be available at the Link.

图像分割多模态频率引导医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。