用小波能量图动态聚焦细节区域,提升2K-4K图像生成质量。
Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
- 基于小波能量图设计频率感知掩码,聚焦潜在空间细节区训练
- 在2K-4K分辨率下显著提升纹理保真度与感知质量,FID得分持续优化
- 无需修改模型结构,训练额外开销为零,适合现有模型升级
高分辨率图像生成仍是生成建模的核心挑战,尤其在计算效率与精细视觉细节保持之间难以平衡。本文提出轻量级训练框架潜空间小波扩散(Latent Wavelet Diffusion, LWD),显著提升2K-4K超高清图像合成中的细节与纹理保真度。LWD引入一种源自小波能量图的新型频率感知掩码策略,动态聚焦潜在空间中细节丰富的区域进行训练。该方法辅以尺度一致的变分自编码器(VAE)目标,确保高谱保真度。其主要优势在于高效性:无需架构修改,推理阶段无额外开销,是扩展现有模型的实用方案。在多个强基线模型上,LWD均持续提升感知质量与FID分数,验证了信号驱动监督在高分辨率生成建模中的有效性和高效性。代码已公开于https://github.com/LuigiSigillo/LatentWaveletDiffusion。
原文摘要 · Abstract (English)
High-resolution image synthesis remains a core challenge in generative modeling, particularly in balancing computational efficiency with the preservation of fine-grained visual detail. We present Latent Wavelet Diffusion (LWD), a lightweight training framework that significantly improves detail and texture fidelity in ultra-high-resolution (2K-4K) image synthesis. LWD introduces a novel, frequency-aware masking strategy derived from wavelet energy maps, which dynamically focuses the training process on detail-rich regions of the latent space. This is complemented by a scale-consistent VAE objective to ensure high spectral fidelity. The primary advantage of our approach is its efficiency: LWD requires no architectural modifications and adds zero additional cost during inference, making it a practical solution for scaling existing models. Across multiple strong baselines, LWD consistently improves perceptual quality and FID scores, demonstrating the power of signal-driven supervision as a principled and efficient path toward high-resolution generative modeling. The code is available at https://github.com/LuigiSigillo/LatentWaveletDiffusion
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。