用DreamBooth微调Stable Diffusion 3,让模型生成更符合提示的图像。
Fine Tuning Text-to-Image Diffusion Models for Correcting Anomalous Images
- 用DreamBooth技术微调Stable Diffusion 3,针对特定提示优化生成效果。
- 在'躺在草地/街道上'提示下,SSIM、PSNR、FID等指标均显著提升。
- 用户偏好调查支持改进效果,适合需要高准确性的图像生成场景。
自生成对抗网络(GAN)和变分自编码器(VAE)问世以来,图像生成模型持续演进,得益于Stable Diffusion和DALL-E等模型的推出,已在艺术、设计和广告等领域实现广泛应用。然而,这些文本到图像模型在某些提示下常生成异常图像。本研究提出一种方法,通过DreamBooth技术对Stable Diffusion 3模型进行微调,以缓解此类问题。针对提示'lying on the grass/street'的实验结果表明,微调后的模型在视觉评估及结构相似性指数(SSIM)、峰值信噪比(PSNR)和弗雷歇起始距离(FID)等指标上均有提升。用户问卷调查也显示,人们对微调后模型生成的图像偏好更高。该研究有望提升文本到图像模型的实际可用性与可靠性。
原文摘要 · Abstract (English)
Since the advent of GANs and VAEs, image generation models have continuously evolved, opening up various real-world applications with the introduction of Stable Diffusion and DALL-E models. These text-to-image models can generate high-quality images for fields such as art, design, and advertising. However, they often produce aberrant images for certain prompts. This study proposes a method to mitigate such issues by fine-tuning the Stable Diffusion 3 model using the DreamBooth technique. Experimental results targeting the prompt "lying on the grass/street" demonstrate that the fine-tuned model shows improved performance in visual evaluation and metrics such as Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Frechet Inception Distance (FID). User surveys also indicated a higher preference for the fine-tuned model. This research is expected to make contributions to enhancing the practicality and reliability of text-to-image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。