arXiv:2607.00987cs.CV2026-07中稿 · ECCV

用扩散模型实现任意缩放的视频超分辨率,解决模糊与帧间抖动问题。

AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution

论文配图:AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
图 1 · 摘自论文原文
  • 分离尺度无关去噪与连续坐标重建,避免重复采样计算。
  • 在16倍缩放下仍保持高频细节,时间稳定性显著优于现有方法。
  • 适合需要高保真动态视频处理的研究者与工业应用。

扩散模型显著推动了视频超分辨率(VSR)发展,但大多局限于固定放大倍数。而基于坐标的任意尺度VSR方法虽具灵活性,但在大缩放时严重过平滑。将生成先验与连续解码结合具有潜力,但受扩散采样随机性影响,常导致严重的时间闪烁。为此,我们提出AVSR-Diff(基于扩散的任意尺度视频超分辨率),一种新型解耦框架:将尺度无关的潜在去噪与连续坐标渲染分离,有效避免计算量巨大的分辨率特异性采样。引入时序门控特征递归(TGFR)模块,提取严格对齐、时序一致的潜在先验。同时设计包含尺度感知傅里叶精修(SAFR)模块的连续视频VAE解码器,动态适配任意目标尺度的频率成分。大量实验表明,AVSR-Diff在多种缩放比例下均能保持高频细节与强时序稳定性,超越当前最先进的任意尺度基线。尤为突出的是,其在原始分辨率上甚至优于近期固定尺度生成模型。

原文摘要 · Abstract (English)

Diffusion models have significantly advanced video super-resolution (VSR) but remain largely constrained to fixed upsampling scales. Conversely, while coordinate-based arbitrary-scale VSR methods offer scale flexibility, they inherently suffer from severe over-smoothing at large scaling factors. Integrating generative priors with continuous decoding is promising but currently hindered by severe temporal flickering caused by the stochasticity of diffusion sampling. To address this, we propose AVSR-Diff (Arbitrary-scale Video Super-Resolution with Diffusion), a novel decoupled framework that separates scale-agnostic latent denoising from continuous coordinate rendering, effectively avoiding computationally heavy resolution-specific sampling. Our approach introduces a Temporally-Gated Feature Recurrence (TGFR) module to extract strictly aligned, temporally consistent latent priors. Furthermore, we design a continuous video VAE decoder incorporating a Scale-Aware Fourier Refinement (SAFR) module to dynamically adapt frequency components to any target scale. Extensive experiments demonstrate that AVSR-Diff consistently preserves high-frequency details and strong temporal stability across various scales, surpassing state-of-the-art arbitrary-scale baselines. Remarkably, our framework outperforms recent fixed-scale generative models even on their native resolution.

视频超分扩散模型任意尺度时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。