用逆演化层抑制生成图像噪声,提升跨域语义分割泛化能力
IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation
- 在扩散模型中加入逆演化层,识别并过滤结构与语义缺陷
- 在多个基准数据集上达到优于现有方法的跨域泛化性能
- 适合研究领域自适应、图像生成质量优化的研究者
领域泛化语义分割(DGSS)旨在利用源域标注数据训练模型,使其在推理时对未见目标域具有鲁棒性。常用方法是使用扩散模型(DMs)生成合成数据以增强源域。然而,由于训练不完善,生成图像常存在结构性或语义缺陷。用此类有缺陷的数据训练分割模型会导致性能下降和错误累积。为此,我们提出在生成过程中引入逆演化层(IELs),其基于拉普拉斯先验突出空间不连续性和语义不一致,实现对不良生成模式的有效过滤。基于此机制,我们提出IELDM,一种增强的基于扩散的数据增强框架,可生成更高质量图像。此外,我们发现IELs的缺陷抑制能力还能通过抑制伪影传播,提升分割网络表现。据此,我们将IELs嵌入DGSS模型解码器,提出IELFormer以强化跨域泛化能力。为进一步提升多尺度语义一致性,IELFormer引入多尺度频率融合(MFF)模块,通过频域分析实现多分辨率特征的结构化融合,改善跨尺度一致性。大量实验表明,本方法在多个基准数据集上显著优于现有方法。
原文摘要 · Abstract (English)
Domain Generalized Semantic Segmentation (DGSS) focuses on training a model using labeled data from a source domain, with the goal of achieving robust generalization to unseen target domains during inference. A common approach to improve generalization is to augment the source domain with synthetic data generated by diffusion models (DMs). However, the generated images often contain structural or semantic defects due to training imperfections. Training segmentation models with such flawed data can lead to performance degradation and error accumulation. To address this issue, we propose to integrate inverse evolution layers (IELs) into the generative process. IELs are designed to highlight spatial discontinuities and semantic inconsistencies using Laplacian-based priors, enabling more effective filtering of undesirable generative patterns. Based on this mechanism, we introduce IELDM, an enhanced diffusion-based data augmentation framework that can produce higher-quality images. Furthermore, we observe that the defect-suppression capability of IELs can also benefit the segmentation network by suppressing artifact propagation. Based on this insight, we embed IELs into the decoder of the DGSS model and propose IELFormer to strengthen generalization capability in cross-domain scenarios. To further strengthen the model's semantic consistency across scales, IELFormer incorporates a multi-scale frequency fusion (MFF) module, which performs frequency-domain analysis to achieve structured integration of multi-resolution features, thereby improving cross-scale coherence. Extensive experiments on benchmark datasets demonstrate that our approach achieves superior generalization performance compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。