arXiv:2412.02929cs.CVcs.AI2024-12

首个可同时生成图像与全景分割图的扩散模型。

Panoptic Diffusion Models: co-generation of images and segmentation maps

  • 通过构建分割布局实现图文对齐,引导生成过程。
  • 支持高分辨率分割图生成,提升背景多样性与类别覆盖。
  • 适用于图像生成与图像到图像转换,适合视觉理解任务。

最近,扩散模型在文本引导和图像条件图像生成方面表现出色。然而,现有模型无法从提示中同时生成图像和全景分割图。融入对形状和场景布局的内在理解,可提升扩散模型的创造力与真实感。为此,我们提出全景扩散模型(PDM),首个设计用于并发生成图像与全景分割图的模型。PDM通过构建分割布局,在生成过程中提供详细内置引导,确保文本提示中提及的类别被包含,并丰富背景中的片段多样性。我们在两种架构上验证了PDM的有效性:统一扩散变压器与带有预训练主干的双流变压器。提出多尺度分块机制以生成高分辨率分割图。此外,当提供真实标签时,PDM可作为文本引导的图像到图像生成模型使用。最后,我们提出一种新型评估指标,证明PDM在隐式场景控制下的图像生成中达到领先水平。

原文摘要 · Abstract (English)

Recently, diffusion models have demonstrated impressive capabilities in text-guided and image-conditioned image generation. However, existing diffusion models cannot simultaneously generate an image and a panoptic segmentation of objects and stuff from the prompt. Incorporating an inherent understanding of shapes and scene layouts can improve the creativity and realism of diffusion models. To address this limitation, we present Panoptic Diffusion Model (PDM), the first model designed to generate both images and panoptic segmentation maps concurrently. PDM bridges the gap between image and text by constructing segmentation layouts that provide detailed, built-in guidance throughout the generation process. This ensures the inclusion of categories mentioned in text prompts and enriches the diversity of segments within the background. We demonstrate the effectiveness of PDM across two architectures: a unified diffusion transformer and a two-stream transformer with a pretrained backbone. We propose a Multi-Scale Patching mechanism to generate high-resolution segmentation maps. Additionally, when ground-truth maps are available, PDM can function as a text-guided image-to-image generation model. Finally, we propose a novel metric for evaluating the quality of generated maps and show that PDM achieves state-of-the-art results in image generation with implicit scene control.

扩散模型全景分割图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。