用全局区域先验实现无训练高清图像生成,兼顾语义一致与创意细节。
Creatively Upscaling Images with Global-Regional Priors
- 通过全局提示和多模态大模型估计区域提示,构建双先验机制。
- 在4096×4096和8192×8192分辨率下生成图像,视觉质量更优。
- 适合需要高质量图像超分且关注局部创意细节的研究者。
当前扩散模型在文本到图像生成方面表现卓越,但仍受限于固定分辨率(如1,024×1,024)。近期方法通过复用预训练模型,利用区域去噪或扩张采样/卷积实现无训练高分辨率生成。然而这些方法难以同时保持全局语义结构并生成富有创意的局部细节。为此,我们提出C-Upscale,一种基于全局-区域先验的无训练图像超分新方法,其先验来自给定全局提示及通过多模态大语言模型(Multimodal LLM)估计的区域提示。技术上,将低分辨率图像的低频成分视为全局结构先验,以保证高分辨率生成中的语义一致性;随后在区域去噪阶段引入区域注意力控制,筛选全局提示与各区域间的交叉注意力,形成区域注意力先验,缓解物体重复问题;进一步,富含描述性细节的估计区域提示作为区域语义先验,激发局部细节的创造性生成。定量与定性评估均表明,我们的C-Upscale可生成4,096×4,096和8,192×8,192等超高清图像,具备更高视觉保真度与更丰富的局部创意细节。
原文摘要 · Abstract (English)
Contemporary diffusion models show remarkable capability in text-to-image generation, while still being limited to restricted resolutions (e.g., 1,024 X 1,024). Recent advances enable tuning-free higher-resolution image generation by recycling pre-trained diffusion models and extending them via regional denoising or dilated sampling/convolutions. However, these models struggle to simultaneously preserve global semantic structure and produce creative regional details in higher-resolution images. To address this, we present C-Upscale, a new recipe of tuning-free image upscaling that pivots on global-regional priors derived from given global prompt and estimated regional prompts via Multimodal LLM. Technically, the low-frequency component of low-resolution image is recognized as global structure prior to encourage global semantic consistency in high-resolution generation. Next, we perform regional attention control to screen cross-attention between global prompt and each region during regional denoising, leading to regional attention prior that alleviates object repetition issue. The estimated regional prompts containing rich descriptive details further act as regional semantic prior to fuel the creativity of regional detail generation. Both quantitative and qualitative evaluations demonstrate that our C-Upscale manages to generate ultra-high-resolution images (e.g., 4,096 X 4,096 and 8,192 X 8,192) with higher visual fidelity and more creative regional details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。