arXiv:2409.00654cs.CV2024-09被引 1

用扩散模型的种子空间实现无配对图像翻译,保持结构语义不变。

Seed-to-Seed: Unpaired Image Translation in Diffusion Seed Space

  • 通过反演扩散模型潜变量生成种子空间,提取语义信息用于翻译。
  • 在复杂汽车场景上优于现有生成模型,结构保留效果更佳。
  • 适合需要高保真图像编辑的自动驾驶视觉任务。

我们提出一种名为Seed-to-Seed Translation(StS)的新方法,结合生成对抗网络(GANs)与扩散模型(DMs),实现无配对图像到图像的翻译。该方法专注于复杂汽车场景的全局翻译,要求严格保持源图像的结构和语义。我们证明,预训练扩散模型中反演潜变量(种子)所构成的种子空间,蕴含可用于判别任务的语义信息,并据此进行图像翻译。方法基于CycleGAN训练一个无配对的种子到种子翻译模型(sts-GAN),将翻译后的种子作为扩散模型采样过程的起点,同时利用ControlNet确保结构一致性。实验表明,该方法在复杂汽车场景的结构保持翻译上表现优异,显著优于现有的基于GAN和扩散模型的方法。除推动汽车场景翻译的最先进水平外,本工作还为利用预训练扩散模型的种子空间进行高效图像编辑提供了新视角。

原文摘要 · Abstract (English)

We introduce Seed-to-Seed Translation (StS), a novel approach that combines GANs and diffusion models (DMs) for unpaired Image-to-Image Translation. Our approach is aimed at global translations of complex automotive scenes, where close adherence to the structure and semantics of the source image is essential. We demonstrate that the semantic information encoded in the space of inverted latents (seeds) of a pretrained DM, dubbed as the seed-space, can be used for discriminative tasks, and leverage this information to perform image-to-image translation. Our method involves training an sts-GAN, an unpaired seed-to-seed translation model, based on CycleGAN. The translated seeds are used as the starting point for the DM's sampling process, while structure preservation is ensured using a ControlNet. We demonstrate the effectiveness of our approach for structure-preserving translation of complex automotive scenes, showcasing superior performance compared to existing GAN-based and diffusion-based methods. In addition to advancing the SoTA in automotive scene translations, our approach offers a fresh perspective on leveraging the semantic information encoded within the seed-space of pretrained DMs for effective image editing and manipulation.

图像翻译扩散模型汽车场景种子空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。