提出AlignVAR框架,提升图像超分的全局一致性与生成效率。
AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution
- 用自适应掩码增强长程依赖,缓解局部注意力偏差。
- 多尺度全图监督,提前暴露误差,稳定渐进式重构过程。
- 比主流扩散模型快10倍、参数少50%,适合高效超分场景。
视觉自回归(VAR)模型因其训练稳定、非迭代推理和高保真合成能力,成为图像生成的新范式,但其在图像超分辨率(ISR)中的应用仍不充分,面临两大挑战:局部注意力偏倚导致空间结构破碎,以及仅依赖残差监督造成误差累积,严重破坏重建图像的全局一致性。为此,我们提出AlignVAR,一种面向ISR的全局一致视觉自回归框架,包含两个核心组件:(1) 空间一致性自回归(SCA),通过自适应掩码重加权注意力至结构相关区域,有效缓解过度局部性,增强长程依赖;(2) 分层一致性约束(HCC),在每尺度引入完整重建监督,提前暴露累积偏差,稳定从粗到精的优化过程。大量实验表明,AlignVAR在结构连贯性和感知保真度上均优于现有生成方法,且推理速度超过扩散模型10倍,参数量减少近50%,确立了高效的新型ISR范式。
原文摘要 · Abstract (English)
Visual autoregressive (VAR) models have recently emerged as a promising alternative for image generation, offering stable training, non-iterative inference, and high-fidelity synthesis through next-scale prediction. This encourages the exploration of VAR for image super-resolution (ISR), yet its application remains underexplored and faces two critical challenges: locality-biased attention, which fragments spatial structures, and residual-only supervision, which accumulates errors across scales, severely compromises global consistency of reconstructed images. To address these issues, we propose AlignVAR, a globally consistent visual autoregressive framework tailored for ISR, featuring two key components: (1) Spatial Consistency Autoregression (SCA), which applies an adaptive mask to reweight attention toward structurally correlated regions, thereby mitigating excessive locality and enhancing long-range dependencies; and (2) Hierarchical Consistency Constraint (HCC), which augments residual learning with full reconstruction supervision at each scale, exposing accumulated deviations early and stabilizing the coarse-to-fine refinement process. Extensive experiments demonstrate that AlignVAR consistently enhances structural coherence and perceptual fidelity over existing generative methods, while delivering over 10x faster inference with nearly 50% fewer parameters than leading diffusion-based approaches, establishing a new paradigm for efficient ISR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。