arXiv:2511.18786cs.CV2025-11

用扩散模型提升视频超分质量,解决运动复杂时的稳定性与结构保真问题。

STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution

  • 分段重建+运动感知设计,应对复杂镜头运动。
  • 利用首帧潜在表示引导生成,提升结构保真度。
  • 适合需要高质量视频重建的科研与工业场景。

我们提出STCDiT,一种基于预训练视频扩散模型的视频超分辨率框架,旨在从退化输入中恢复结构准确且时间稳定的视频,即使在复杂相机运动下也能保持效果。主要挑战在于重建过程中的时间稳定性与生成过程中的结构保真性。为此,我们首先设计了一种运动感知的VAE重建方法,对视频进行分段重建,每段具有均匀运动特征,有效处理复杂运动视频。此外,我们发现每个片段中由VAE编码器提取的首帧潜在表示(称为锚帧潜在)不受时间压缩影响,保留了比后续帧更丰富的空间结构信息。因此,我们进一步提出锚帧引导方法,利用锚帧的结构信息约束生成过程,提升视频特征的结构保真度。结合两项设计,使视频扩散模型实现高质量视频超分辨率。大量实验表明,STCDiT在结构保真度和时间一致性上均优于现有最先进方法。

原文摘要 · Abstract (English)

We present STCDiT, a video super-resolution framework built upon a pre-trained video diffusion model, aiming to restore structurally faithful and temporally stable videos from degraded inputs, even under complex camera motions. The main challenges lie in maintaining temporal stability during reconstruction and preserving structural fidelity during generation. To address these challenges, we first develop a motion-aware VAE reconstruction method that performs segment-wise reconstruction, with each segment clip exhibiting uniform motion characteristic, thereby effectively handling videos with complex camera motions. Moreover, we observe that the first-frame latent extracted by the VAE encoder in each clip, termed the anchor-frame latent, remains unaffected by temporal compression and retains richer spatial structural information than subsequent frame latents. We further develop an anchor-frame guidance approach that leverages structural information from anchor frames to constrain the generation process and improve structural fidelity of video features. Coupling these two designs enables the video diffusion model to achieve high-quality video super-resolution. Extensive experiments show that STCDiT outperforms state-of-the-art methods in terms of structural fidelity and temporal consistency.

视频超分扩散模型时空一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。