针对生成式超分辨率中的纹理建模误差,提出新编码与重建感知预测方法。
Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- 仅用码本建模缺失纹理,减少编码误差
- 通过图像级监督训练索引预测器,提升重建精度
- 轻量级设计,实现逼真超分效果
基于向量量化(VQ)的模型在视觉先验建模方面展现出强大潜力。然而,现有VQ方法仅通过最近码本项编码视觉特征,并以码本级别监督训练索引预测器。由于视觉信号丰富,VQ编码常导致较大量化误差。此外,以码本级监督训练预测器无法考虑最终重建误差,造成先验建模精度不足。本文针对上述问题,提出纹理向量量化(Texture VQ)与重建感知预测(Reconstruction Aware Prediction)策略。纹理向量量化利用超分辨率任务特性,仅对缺失纹理引入码本建模;重建感知预测则利用直通估计器,以图像级监督直接训练索引预测器。所提出的生成式超分模型(TVQ&RAP)在计算开销小的前提下,可生成逼真的超分结果。
原文摘要 · Abstract (English)
Vector-quantized based models have recently demonstrated strong potential for visual prior modeling. However, existing VQ-based methods simply encode visual features with nearest codebook items and train index predictor with code-level supervision. Due to the richness of visual signal, VQ encoding often leads to large quantization error. Furthermore, training predictor with code-level supervision can not take the final reconstruction errors into consideration, result in sub-optimal prior modeling accuracy. In this paper we address the above two issues and propose a Texture Vector-Quantization and a Reconstruction Aware Prediction strategy. The texture vector-quantization strategy leverages the task character of super-resolution and only introduce codebook to model the prior of missing textures. While the reconstruction aware prediction strategy makes use of the straight-through estimator to directly train index predictor with image-level supervision. Our proposed generative SR model (TVQ&RAP) is able to deliver photo-realistic SR results with small computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。