用小波系数构建多尺度潜在空间,提升图像生成清晰度。
Wavelet-based Variational Autoencoders for High-Resolution Image Generation
- 用哈尔小波分解图像,分层编码细节与近似系数
- 生成图像在CIFAR-10上视觉质量显著优于传统VAE
- 适合追求高分辨率细节生成的图像建模任务
变分自编码器(VAEs)能学习紧凑的潜在表示,但传统VAE因假设各向同性高斯潜空间且难以捕捉高频细节,常生成模糊图像。本文提出一种基于小波的新方法(Wavelet-VAE),将潜空间构建为多尺度哈尔小波系数。我们系统地将图像特征编码为多尺度细节与近似系数,并引入可学习噪声参数以保持随机性。深入探讨了重参数化技巧的重构、KL散度项的处理,以及小波稀疏性原则在训练目标中的融合。在CIFAR-10及其他高分辨率数据集上的实验表明,Wavelet-VAE在视觉保真度和高分辨率细节恢复方面优于传统VAE。最后讨论了该方法的优势、潜在局限及未来研究方向。
原文摘要 · Abstract (English)
Variational Autoencoders (VAEs) are powerful generative models capable of learning compact latent representations. However, conventional VAEs often generate relatively blurry images due to their assumption of an isotropic Gaussian latent space and constraints in capturing high-frequency details. In this paper, we explore a novel wavelet-based approach (Wavelet-VAE) in which the latent space is constructed using multi-scale Haar wavelet coefficients. We propose a comprehensive method to encode the image features into multi-scale detail and approximation coefficients and introduce a learnable noise parameter to maintain stochasticity. We thoroughly discuss how to reformulate the reparameterization trick, address the KL divergence term, and integrate wavelet sparsity principles into the training objective. Our experimental evaluation on CIFAR-10 and other high-resolution datasets demonstrates that the Wavelet-VAE improves visual fidelity and recovers higher-resolution details compared to conventional VAEs. We conclude with a discussion of advantages, potential limitations, and future research directions for wavelet-based generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。