用扩散模型生成差异改进不精确分割,无需复杂结构
G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

- 通过原始图与掩码条件生成图的差异捕捉语义对应关系
- 实现从粗到细的分割优化,提升前景概率估计精度
- 插件式设计适配现有模型,适合图像分割研究者
本文研究如何利用大规模文本到图像扩散模型解决不精确分割(IS)难题。不同于依赖判别模型或内部注意力机制提取密集视觉表征的传统方法,本方法聚焦于Stable Diffusion(SD)固有的生成先验。具体而言,通过分析原始图像与掩码条件生成图像之间的模式差异,建立语义对应关系并更新前景概率,实现粗到细的分割精化。大量定量与定性实验验证了该方法的有效性与优越性,凸显利用生成差异建模密集表征的潜力,推动生成式方法在判别任务中的进一步探索。
原文摘要 · Abstract (English)
This paper considers the problem of utilizing a large-scale text-to-image diffusion model to tackle the challenging Inexact Segmentation (IS) task. Unlike traditional approaches that rely heavily on discriminative-model-based paradigms or dense visual representations derived from internal attention mechanisms, our method focuses on the intrinsic generative priors in Stable Diffusion~(SD). Specifically, we exploit the pattern discrepancies between original images and mask-conditional generated images to facilitate a coarse-to-fine segmentation refinement by establishing a semantic correspondence alignment and updating the foreground probability. Comprehensive quantitative and qualitative experiments validate the effectiveness and superiority of our plug-and-play design, underscoring the potential of leveraging generation discrepancies to model dense representations and encouraging further exploration of generative approaches for solving discriminative tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。