arXiv:2503.12507cs.CV2025-03CVPR被引 6

让分割模型在模糊低质图像上仍能准确分割,提升真实场景适用性。

Segment Any-Quality Images with Generative Latent Space Enhancement

  • 在SAM的潜在空间中做生成式修复,重建高质量图像表征。
  • 在复杂退化条件下分割性能显著提升,且对未见过的退化也有效。
  • 仅需少量新增参数,可适配现有SAM模型,适合部署落地。

尽管表现优异,现有分割任意图像模型(SAM)在严重退化的低质量图像上性能显著下降,限制了其在真实场景中的应用。为此,本文提出GleSAM,通过生成式潜在空间增强提升模型对低质图像的鲁棒性,实现跨图像质量的泛化能力。具体地,将潜在扩散思想引入SAM分割框架,在SAM的潜在空间中执行生成式扩散过程,以重建高质量表征并提升分割效果。同时,设计两项技术提升预训练扩散模型与分割框架的兼容性。该方法仅需极少量新增可学习参数,即可应用于预训练的SAM和SAM2,实现高效优化。此外,构建了包含多种退化类型与程度的LQSeg数据集,用于模型训练与评估。大量实验表明,GleSAM在复杂退化条件下显著提升分割鲁棒性,同时保持对清晰图像的良好泛化能力。更重要的是,其在未见退化类型上也表现良好,验证了方法与数据集的通用性。

原文摘要 · Abstract (English)

Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose GleSAM, which utilizes Generative Latent space Enhancement to boost robustness on low-quality images, thus enabling generalization across various image qualities. Specifically, we adapt the concept of latent diffusion to SAM-based segmentation frameworks and perform the generative diffusion process in the latent space of SAM to reconstruct high-quality representation, thereby improving segmentation. Additionally, we introduce two techniques to improve compatibility between the pre-trained diffusion model and the segmentation framework. Our method can be applied to pre-trained SAM and SAM2 with only minimal additional learnable parameters, allowing for efficient optimization. We also construct the LQSeg dataset with a greater diversity of degradation types and levels for training and evaluating the model. Extensive experiments demonstrate that GleSAM significantly improves segmentation robustness on complex degradations while maintaining generalization to clear images. Furthermore, GleSAM also performs well on unseen degradations, underscoring the versatility of our approach and dataset.

图像分割生成模型鲁棒性潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。