arXiv:2412.09626cs.CV2024-12ICCV被引 31

无需微调,让扩散模型生成8K超清图像

FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion

  • 通过多尺度信息融合提取所需频段,实现无微调高分辨率生成
  • 首次实现8K文本到图像生成,显著减少重复图案
  • 适合追求极致画质的图像/视频生成研究者

视觉扩散模型取得显著进展,但受限于高质量高分辨率数据稀缺和计算资源不足,通常在有限分辨率下训练,制约了生成更高分辨率、高保真度图像或视频的能力。近期方法尝试无微调策略以释放预训练模型的潜在高分辨率生成能力,但仍易产生低质量内容及重复模式。核心问题在于:当生成超出训练分辨率的视觉内容时,高频信息不可避免增加,导致累积误差引发重复模式。为此,我们提出FreeScale——一种无微调的推理范式,通过多尺度处理与频段选择性融合,实现高分辨率视觉生成。大量实验验证其在图像与视频模型上的优越性。尤为关键的是,相比此前最优方法,FreeScale首次实现8K文本到图像生成。

原文摘要 · Abstract (English)

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity images or videos at higher resolutions. Recent efforts have explored tuning-free strategies to exhibit the untapped potential higher-resolution visual generation of pre-trained models. However, these methods are still prone to producing low-quality visual content with repetitive patterns. The key obstacle lies in the inevitable increase in high-frequency information when the model generates visual content exceeding its training resolution, leading to undesirable repetitive patterns deriving from the accumulated errors. To tackle this challenge, we propose FreeScale, a tuning-free inference paradigm to enable higher-resolution visual generation via scale fusion. Specifically, FreeScale processes information from different receptive scales and then fuses it by extracting desired frequency components. Extensive experiments validate the superiority of our paradigm in extending the capabilities of higher-resolution visual generation for both image and video models. Notably, compared with previous best-performing methods, FreeScale unlocks the 8k-resolution text-to-image generation for the first time.

扩散模型超分辨率图像生成无微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。