arXiv:2411.15925cs.CVcs.AI2024-11

通过重排图像区域,实现任意主题的新图像生成。

Making Images from Images: Interleaving Denoising and Transformation

  • 交替执行图像去噪与能量优化,实现图像变换学习。
  • 区域越多结果越优,突破传统数量限制。
  • 支持像素与潜在空间,可融合多源图像创意生成。

仅通过重新排列图像的区域,即可生成任意主题的新图像。区域定义由用户自定,包括规则或不规则块、同心圆甚至单个像素。该方法扩展并改进了近期关于光学幻觉生成的工作,同时学习图像内容及参数化变换以相互转换目标图像。通过学习图像变换,可预先指定任意源图像;任何现有图像(如《蒙娜丽莎》)均可转化为新主题。我们将此过程建模为约束优化问题,通过交替执行图像扩散与能量最小化步骤求解。与以往方法不同,增加区域数量反而使问题更易解决且提升效果。我们在像素空间和潜在空间中均验证了该方法的有效性。还提出了创造性扩展,如使用源图像无限副本或多源图像融合。

原文摘要 · Abstract (English)

Simply by rearranging the regions of an image, we can create a new image of any subject matter. The definition of regions is user definable, ranging from regularly and irregularly-shaped blocks, concentric rings, or even individual pixels. Our method extends and improves recent work in the generation of optical illusions by simultaneously learning not only the content of the images, but also the parameterized transformations required to transform the desired images into each other. By learning the image transforms, we allow any source image to be pre-specified; any existing image (e.g. the Mona Lisa) can be transformed to a novel subject. We formulate this process as a constrained optimization problem and address it through interleaving the steps of image diffusion with an energy minimization step. Unlike previous methods, increasing the number of regions actually makes the problem easier and improves results. We demonstrate our approach in both pixel and latent spaces. Creative extensions, such as using infinite copies of the source image and employing multiple source images, are also given.

图像生成扩散模型图像变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。