构建10万张超高清图像数据集,提升文本生成超清图的细节质量。
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
- 构建10万张超高清图像数据集,支持高精度图像生成训练。
- 提出频率感知微调方法,在3000+分辨率下显著提升细节还原能力。
- 适合关注超清图像生成、细节优化的研究者与开发者。
超高清(UHR)文本到图像生成已取得显著进展,但仍面临两大挑战:一是缺乏大规模高质量的UHR T2I数据集;二是未针对超高清场景下的精细细节合成设计专用训练策略。为解决第一项挑战,我们提出 extbf{UltraHR-100K},一个包含10万张高保真超高清图像及其丰富描述的高质量数据集,每张图像分辨率均超过3000,且在细节丰富度、内容复杂度和美学质量上经过严格筛选。为应对第二项挑战,我们提出一种频率感知后训练方法,通过两项创新实现:(i) extit{面向细节的时间步采样(DOTS)},聚焦于对细节重建至关重要的去噪步骤;(ii) extit{软权重频域正则化(SWFR)},利用离散傅里叶变换(DFT)对频域成分进行软约束,强化高频细节保留。在所提出的UltraHR-eval4K基准测试上,实验表明该方法显著提升了超高清图像生成的精细细节质量与整体保真度。代码已开源。
原文摘要 · Abstract (English)
Ultra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce \textbf{UltraHR-100K}, a high-quality dataset of 100K UHR images with rich captions, offering diverse content and strong visual fidelity. Each image exceeds 3K resolution and is rigorously curated based on detail richness, content complexity, and aesthetic quality. To tackle the second challenge, we propose a frequency-aware post-training method that enhances fine-detail generation in T2I diffusion models. Specifically, we design (i) \textit{Detail-Oriented Timestep Sampling (DOTS)} to focus learning on detail-critical denoising steps, and (ii) \textit{Soft-Weighting Frequency Regularization (SWFR)}, which leverages Discrete Fourier Transform (DFT) to softly constrain frequency components, encouraging high-frequency detail preservation. Extensive experiments on our proposed UltraHR-eval4K benchmarks demonstrate that our approach significantly improves the fine-grained detail quality and overall fidelity of UHR image generation. The code is available at \href{https://github.com/NJU-PCALab/UltraHR-100k}{here}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。