通过可逆变换与语义先验,实现高倍率图像缩放时结构一致与细节真实。
Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors

- 基于扩散模型设计可逆变换,在潜在空间实现双向缩放。
- 引入自适应高频词典补偿量化损失,提升细节还原能力。
- 轻量级语义嵌入器增强生成一致性,适合高倍率图像修复场景。
现有高倍率图像缩放方法在16倍及以上缩放因子下,因低到高分辨率映射的病态性,难以保持语义一致结构并生成真实细节。为此,我们提出基于扩散模型的FaithEIR框架。受奇异值分解启发,设计可学习的可逆变换,实现潜在空间中的可逆下采样与上采样。为补偿量化带来的信息丢失,提出自适应细节先验——一个捕捉训练数据中常见结构统计平均的高频词典。同时,设计轻量级像素语义嵌入器,为预训练扩散模型提供语义条件。大量实验表明,FaithEIR持续优于当前最优方法,在重建保真度和感知质量上均表现更优。代码、模型权重及详细结果已开源:https://github.com/cshw2021/FaithEIR。
原文摘要 · Abstract (English)
Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-posed nature of low- to high-resolution mapping under scaling factors of $16\times$ or higher. To alleviate the above problems, we propose FaithEIR, a diffusion-based framework for extreme image rescaling. Inspired by singular value decomposition, we develop learnable reversible transformation that enables invertible downscaling and upscaling in the latent space. To compensate for information loss due to quantization, we propose an adaptive detail prior, a high-frequency dictionary that captures the empirical average of commonly occurring structures in the training data. Finally, we design a lightweight pixel semantic embedder to provide semantic conditioning for the pretrained diffusion model. We present extensive experimental results demonstrating that our FaithEIR consistently outperforms state-of-the-art methods, achieving superior reconstruction fidelity and perceptual quality. Our code, model weights, and detailed results are released at https://github.com/cshw2021/FaithEIR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。