让扩散模型无须训练就能生成8K超分辨率图像。
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
- 用多路径扩散+局部退化提示,实现高分辨率生成。
- 无需额外训练,直接生成2K/4K/8K图像。
- 适合需要超高清图像生成的科研与工业场景。
大规模预训练的文本到图像(T2I)扩散模型在图像生成任务中广受欢迎,并展现出意外的超分辨率(SR)潜力。然而,大多数现有T2I扩散模型的训练分辨率上限为512x512,如何突破此限制成为图像超分辨率领域亟待解决但尚未攻克的难题。本文提出一种新方法,首次实现无需任何额外训练即可生成2K、4K甚至8K图像。该方法结合多扩散(MultiDiffusion)技术,通过多条扩散路径分布生成以保证大尺度下的全局一致性;同时引入局部退化感知提示提取机制,引导T2I模型根据低分辨率输入重建精细局部结构。这些创新使T2I扩散模型可无限制地应用于图像超分辨率任务,显著提升生成分辨率。
原文摘要 · Abstract (English)
Large-scale, pre-trained Text-to-Image (T2I) diffusion models have gained significant popularity in image generation tasks and have shown unexpected potential in image Super-Resolution (SR). However, most existing T2I diffusion models are trained with a resolution limit of 512x512, making scaling beyond this resolution an unresolved but necessary challenge for image SR. In this work, we introduce a novel approach that, for the first time, enables these models to generate 2K, 4K, and even 8K images without any additional training. Our method leverages MultiDiffusion, which distributes the generation across multiple diffusion paths to ensure global coherence at larger scales, and local degradation-aware prompt extraction, which guides the T2I model to reconstruct fine local structures according to its low-resolution input. These innovations unlock higher resolutions, allowing T2I diffusion models to be applied to image SR tasks without limitation on resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。