arXiv:2506.20452cs.CVcs.LG2025-06SIGGRAPH被引 7

无需训练即可生成超高清图像,显著减少伪影,提升细节质量。

HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling

  • 通过两阶段流程:先生成基础图像,再用小波域增强器修复细节。
  • 在Stable Diffusion XL上测试,80%以上用户更偏好其生成结果。
  • 不需重新训练或改架构,适合追求高质量图像的创作者使用。

扩散模型已成为图像生成的主流方法,具备出色的逼真度和多样性。然而,在高分辨率下训练扩散模型仍计算成本高昂,现有零样本生成技术在超出训练分辨率时往往产生伪影,如物体重复和空间不一致。本文提出HiWave,一种无需训练的零样本方法,利用预训练扩散模型显著提升超高清图像合成的视觉保真度与结构一致性。该方法采用两阶段流程:首先从预训练模型生成基础图像,随后进行分块DDIM反演并引入新型小波域细节增强模块。具体而言,先通过反演方法获取保持全局一致性的初始噪声向量;采样过程中,小波域增强器保留基础图像的低频成分以保证结构连贯性,同时有选择地引导高频成分以丰富细节与纹理。在Stable Diffusion XL上的大量实验表明,HiWave有效缓解了以往方法中的常见视觉伪影,显著提升感知质量。用户研究证实,其在超过80%的对比中胜过当前最优方法,证明其在无需重训练或架构修改的前提下,可高效实现高质量超高清图像合成。

原文摘要 · Abstract (English)

Diffusion models have emerged as the leading approach for image synthesis, demonstrating exceptional photorealism and diversity. However, training diffusion models at high resolutions remains computationally prohibitive, and existing zero-shot generation techniques for synthesizing images beyond training resolutions often produce artifacts, including object duplication and spatial incoherence. In this paper, we introduce HiWave, a training-free, zero-shot approach that substantially enhances visual fidelity and structural coherence in ultra-high-resolution image synthesis using pretrained diffusion models. Our method employs a two-stage pipeline: generating a base image from the pretrained model followed by a patch-wise DDIM inversion step and a novel wavelet-based detail enhancer module. Specifically, we first utilize inversion methods to derive initial noise vectors that preserve global coherence from the base image. Subsequently, during sampling, our wavelet-domain detail enhancer retains low-frequency components from the base image to ensure structural consistency, while selectively guiding high-frequency components to enrich fine details and textures. Extensive evaluations using Stable Diffusion XL demonstrate that HiWave effectively mitigates common visual artifacts seen in prior methods, achieving superior perceptual quality. A user study confirmed HiWave's performance, where it was preferred over the state-of-the-art alternative in more than 80% of comparisons, highlighting its effectiveness for high-quality, ultra-high-resolution image synthesis without requiring retraining or architectural modifications.

图像生成扩散模型小波变换超分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。