用小波微调实现4K图像直接生成,解决超分辨率图像合成难题。
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
- 基于小波变换的微调方法,直接训练4K级图像生成。
- 在4K图像上达到高保真度与提示词一致性,性能优于现有模型。
- 构建首个公开4K图像合成基准,支持细节与整体质量评估。
本文提出Diffusion-4K,一种基于文本到图像扩散模型的直接超高清图像生成框架。核心贡献包括:(1) 构建Aesthetic-4K基准数据集,包含经GPT-4o生成标注的高质量4K图像,引入GLCM评分和压缩比指标评估细粒度细节,并结合FID、美学评分与CLIPScore进行综合评估;(2) 提出基于小波的微调方法,适用于多种潜空间扩散模型,在真实4K图像上实现直接训练,显著提升细节还原能力。实验表明,该方法在现代大规模扩散模型(如SD3-2B和Flux-12B)驱动下,显著提升了超高清图像生成质量与提示遵循性,验证了其在4K图像合成中的优越性。
原文摘要 · Abstract (English)
In this paper, we present Diffusion-4K, a novel framework for direct ultra-high-resolution image synthesis using text-to-image diffusion models. The core advancements include: (1) Aesthetic-4K Benchmark: addressing the absence of a publicly available 4K image synthesis dataset, we construct Aesthetic-4K, a comprehensive benchmark for ultra-high-resolution image generation. We curated a high-quality 4K dataset with carefully selected images and captions generated by GPT-4o. Additionally, we introduce GLCM Score and Compression Ratio metrics to evaluate fine details, combined with holistic measures such as FID, Aesthetics and CLIPScore for a comprehensive assessment of ultra-high-resolution images. (2) Wavelet-based Fine-tuning: we propose a wavelet-based fine-tuning approach for direct training with photorealistic 4K images, applicable to various latent diffusion models, demonstrating its effectiveness in synthesizing highly detailed 4K images. Consequently, Diffusion-4K achieves impressive performance in high-quality image synthesis and text prompt adherence, especially when powered by modern large-scale diffusion models (e.g., SD3-2B and Flux-12B). Extensive experimental results from our benchmark demonstrate the superiority of Diffusion-4K in ultra-high-resolution image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。