arXiv:2410.02369cs.CV2024-10NeurIPS被引 29

用扩散模型提升少样本语义分割性能,实现高效图像理解。

Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation

论文配图:Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
图 1 · 摘自论文原文
  • 在自注意力机制中设计键值融合方法,增强查询图与支持图交互。
  • 在多个设置下超越现有最先进模型,显著提升分割精度。
  • 适合研究通用分割模型或扩散模型应用的学者参考。

扩散模型不仅在图像生成领域取得显著成果,还展现出利用无标签数据进行有效预训练的潜力。基于其在语义对应和开放词汇分割中的表现,本文首次探索将潜在扩散模型应用于少样本语义分割。受大型语言模型上下文学习能力启发,少样本语义分割已演变为上下文分割任务,成为评估通用分割模型的关键指标。本文聚焦于少样本分割,为基于扩散模型的通用分割模型发展奠定基础。通过分析查询图像与支持图像的交互机制,提出一种自注意力框架内的键值融合方法;进一步优化支持掩码信息注入,并重新评估查询掩码的合理监督方式。基于上述分析,提出名为DiffewS的简单高效框架,最大程度保留原始潜在扩散模型的生成结构并有效利用预训练先验。实验表明,该方法在多种设置下显著优于现有最先进模型。

原文摘要 · Abstract (English)

The Diffusion Model has not only garnered noteworthy achievements in the realm of image generation but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data. Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation. Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models. In this context, we concentrate on Few-shot Semantic Segmentation, establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework. Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask. Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings.

扩散模型少样本分割语义分割自注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。