arXiv:2409.17565cs.CVcs.AI2024-09被引 5

让扩散模型在像素空间优化,显著提升细节生成质量

Pixel-Space Post-Training of Latent Diffusion Models

  • 后训练阶段引入像素空间监督,弥补潜空间细节损失
  • 在DiT和U-Net模型上,视觉质量和缺陷指标大幅提升
  • 适合关注图像细节与真实感的生成应用开发者

近年来,潜空间扩散模型(LDMs)在图像生成领域取得显著进展。其优势在于在压缩的潜空间中操作,实现更高效的训练与部署。然而,仍存在挑战:LDMs常在高频细节和复杂构图上表现不佳。我们推测,这源于预训练和后训练均在潜空间进行,而潜空间分辨率通常比输出图像低8×8倍。为此,我们提出在后训练过程中加入像素空间监督,以更好保留高频细节。实验表明,在最先进的DiT Transformer和U-Net扩散模型上,加入像素空间目标可大幅改善监督式质量微调和基于偏好后训练的效果,视觉质量与视觉缺陷指标均有显著提升,同时保持原有的文本对齐能力。

原文摘要 · Abstract (English)

Latent diffusion models (LDMs) have made significant advancements in the field of image generation in recent years. One major advantage of LDMs is their ability to operate in a compressed latent space, allowing for more efficient training and deployment. However, despite these advantages, challenges with LDMs still remain. For example, it has been observed that LDMs often generate high-frequency details and complex compositions imperfectly. We hypothesize that one reason for these flaws is due to the fact that all pre- and post-training of LDMs are done in latent space, which is typically $8 \times 8$ lower spatial-resolution than the output images. To address this issue, we propose adding pixel-space supervision in the post-training process to better preserve high-frequency details. Experimentally, we show that adding a pixel-space objective significantly improves both supervised quality fine-tuning and preference-based post-training by a large margin on a state-of-the-art DiT transformer and U-Net diffusion models in both visual quality and visual flaw metrics, while maintaining the same text alignment quality.

扩散模型图像生成细节优化后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。