将流模型引入潜在空间,提升人脸修复的感知质量与保真度平衡。
Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration
- 在变分自编码器的潜在空间重构PMRF,更好匹配人类视觉感知。
- 相比原始PMRF,FID指标提升5.79倍速度,实现更优感知-失真权衡。
- 适用于追求高保真、高质量人脸修复的研究者与工程应用。
感知-失真权衡理论指出,人脸修复需在感知质量与保真度间取得平衡。现有基于流模型的后验均值修正流(PMRF)虽有效,但其像素空间建模难以对齐人类感知——即人眼区分两图像分布的能力。本文提出潜空间PMRF(Latent-PMRF),将PMRF重构于变分自编码器(VAE)的潜在空间,优化过程更契合人类感知。通过在最小失真估计的潜在表示上定义源分布,将最小失真约束为VAE的重建误差。实验表明,所提VAE在重建与修复任务中显著优于现有方法。在盲人脸修复任务中,Latent-PMRF展现出更优的感知-失真权衡,并实现5.79倍于PMRF的收敛效率提升(以FID衡量)。代码将开源。
原文摘要 · Abstract (English)
The Perception-Distortion tradeoff (PD-tradeoff) theory suggests that face restoration algorithms must balance perceptual quality and fidelity. To achieve minimal distortion while maintaining perfect perceptual quality, Posterior-Mean Rectified Flow (PMRF) proposes a flow based approach where source distribution is minimum distortion estimations. Although PMRF is shown to be effective, its pixel-space modeling approach limits its ability to align with human perception, where human perception is defined as how humans distinguish between two image distributions. In this work, we propose Latent-PMRF, which reformulates PMRF in the latent space of a variational autoencoder (VAE), facilitating better alignment with human perception during optimization. By defining the source distribution on latent representations of minimum distortion estimation, we bound the minimum distortion by the VAE's reconstruction error. Moreover, we reveal the design of VAE is crucial, and our proposed VAE significantly outperforms existing VAEs in both reconstruction and restoration. Extensive experiments on blind face restoration demonstrate the superiority of Latent-PMRF, offering an improved PD-tradeoff compared to existing methods, along with remarkable convergence efficiency, achieving a 5.79X speedup over PMRF in terms of FID. Our code will be available as open-source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。