构建4K图像生成数据集与方法,实现高分辨率图像高效合成。
Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- 提出Diffusion-4K框架,融合SC-VAE与小波微调,实现4K图像直接生成。
- 在Flux-12B等大模型下生成细节丰富、纹理真实的4K图像。
- 设计GLCM Score与压缩比等新指标,全面评估超高清图像质量。
超高清图像合成具有重要潜力,但因缺乏标准基准和计算限制而研究不足。本文构建了精心筛选的Aesthetic-4K数据集,包含用于训练和评估的专用子集,包含高质量4K图像及GPT-4o生成的描述性标题。提出Diffusion-4K框架,通过尺度一致变分自编码器(SC-VAE)和基于小波的潜在微调(WLF),实现高效视觉标记压缩并捕捉超高清图像中的精细细节,支持直接使用真实感4K数据进行训练。该方法适用于多种潜在扩散模型,在生成高度细节化的4K图像方面表现优异。此外,提出新评价指标GLCM Score与压缩比,结合FID、美学评分与CLIPScore等整体指标,实现多维度评估。Diffusion-4K在先进大规模扩散模型(如Flux-12B)驱动下取得显著性能。源代码公开于https://github.com/zhang0jhon/diffusion-4k。
原文摘要 · Abstract (English)
Ultra-high-resolution image synthesis holds significant potential, yet remains an underexplored challenge due to the absence of standardized benchmarks and computational constraints. In this paper, we establish Aesthetic-4K, a meticulously curated dataset containing dedicated training and evaluation subsets specifically designed for comprehensive research on ultra-high-resolution image synthesis. This dataset consists of high-quality 4K images accompanied by descriptive captions generated by GPT-4o. Furthermore, we propose Diffusion-4K, an innovative framework for the direct generation of ultra-high-resolution images. Our approach incorporates the Scale Consistent Variational Auto-Encoder (SC-VAE) and Wavelet-based Latent Fine-tuning (WLF), which are designed for efficient visual token compression and the capture of intricate details in ultra-high-resolution images, thereby facilitating direct training with photorealistic 4K data. This method is applicable to various latent diffusion models and demonstrates its efficacy in synthesizing highly detailed 4K images. Additionally, we propose novel metrics, namely the GLCM Score and Compression Ratio, to assess the texture richness and fine details in local patches, in conjunction with holistic measures such as FID, Aesthetics, and CLIPScore, enabling a thorough and multifaceted evaluation of ultra-high-resolution image synthesis. Consequently, Diffusion-4K achieves impressive performance in ultra-high-resolution image synthesis, particularly when powered by state-of-the-art large-scale diffusion models (eg, Flux-12B). The source code is publicly available at https://github.com/zhang0jhon/diffusion-4k.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。