分层编码器解耦图像结构与纹理,提升真实场景超分辨率质量
HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution

- 分通道构建层级代码本,分离低频结构与高频纹理
- 在Real-ISR基准上实现最优感知质量与重建精度
- 适合追求高保真图像重建的研究者与开发者
向量量化(VQ)生成模型在真实世界图像超分辨率(Real-ISR)中表现优异。然而,现有方法通常依赖单一潜空间,将低频结构与高频纹理混合编码,导致单个代码本需处理组合复杂的结构-纹理配对,限制表征能力并影响代码本利用率。为此,我们提出HiTokSR,一种分粗到细的分层令牌预测框架。通过沿通道维度将潜空间划分为频率感知组,每组使用独立子代码本进行量化,实现全局结构与细节的解耦,增强组合表达力,同时避免高维最近邻查找带来的优化不稳定性。为提升语义一致性,生成器通过自适应特征调制、多尺度类别令牌及表示对齐损失引入视觉基础模型先验。此外,在解码器微调阶段采用索引级扰动策略,缓解离散令牌预测中的训练-测试差异。在真实世界基准上的大量实验表明,HiTokSR在感知质量和重建保真度方面均达到当前最优水平。
原文摘要 · Abstract (English)
Vector-quantized (VQ) generative models have shown promising results in real-world image super-resolution (Real-ISR). However, existing methods typically rely on a monolithic latent space that entangles low-frequency structures with high-frequency textures. This entanglement forces a single codebook to capture a combinatorially complex set of structure-texture pairings, which constrains representational capacity and limits codebook utilization. To address this issue, we present HiTokSR, a hierarchical token prediction framework. Instead of using a single codebook, HiTokSR partitions the latent space along the channel dimension into frequency-aware groups, quantizing each with an independent sub-codebook. This coarse-to-fine design disentangles global structures from fine details, enhancing combinatorial expressiveness while circumventing the optimization instability of high-dimensional nearest-neighbor lookups. To further improve semantic consistency, our generator integrates priors from a vision foundation model via adaptive feature modulation, multi-scale class tokens, and a representation alignment loss. Additionally, we introduce an index-level perturbation strategy during decoder fine-tuning to bridge the train-test discrepancy in discrete token prediction. Extensive experiments on real-world benchmarks demonstrate that HiTokSR achieves state-of-the-art performance in both perceptual quality and reconstruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。