arXiv:2411.07126cs.CVcs.LG2024-11被引 25

用分频衰减扩散模型实现像素级精准的高质量图像生成

Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

  • 采用分层像素空间扩散,按频段不同速率衰减图像信号
  • 可生成4K超清图像,支持文本生成、图像超分等多场景应用
  • 适合追求像素级精度的图像生成任务,如高保真内容创作

我们提出Edify Image,一类能够以像素级精度生成逼真图像内容的扩散模型。该模型采用级联的像素空间扩散架构,通过一种新颖的拉普拉斯扩散过程训练,使图像在不同频率带上的信号以不同速率衰减。Edify Image支持多种应用场景,包括文本到图像生成、4K图像上采样、ControlNets控制生成、360° HDR全景图生成以及图像定制化微调。

原文摘要 · Abstract (English)

We introduce Edify Image, a family of diffusion models capable of generating photorealistic image content with pixel-perfect accuracy. Edify Image utilizes cascaded pixel-space diffusion models trained using a novel Laplacian diffusion process, in which image signals at different frequency bands are attenuated at varying rates. Edify Image supports a wide range of applications, including text-to-image synthesis, 4K upsampling, ControlNets, 360 HDR panorama generation, and finetuning for image customization.

图像生成扩散模型像素精度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。