arXiv:2503.03355cs.CVcs.LG2025-03被引 2

用扩散模型实现无需运动对齐的视频超分辨率,提升生成质量。

Rethinking Video Super-Resolution: Towards Diffusion-Based Methods without Motion Alignment

  • 基于扩散后验采样框架,使用无条件潜空间视频扩散变压器。
  • 在合成与真实数据集上均实现高质量超分辨率,无需显式光流估计。
  • 单个模型可适配多种采样条件,无需重新训练,通用性强。

本文重新思考视频超分辨率方法,提出一种基于扩散后验采样框架的方法,结合在潜空间运行的无条件视频扩散变换器。该视频生成模型作为时空模型,认为一个能学习真实世界物理规律的强大模型,可自然掌握各种运动模式作为先验知识,从而无需显式估计光流或运动参数进行像素对齐。此外,所提视频扩散变换器模型仅需单一实例即可适应不同采样条件而无需重训练。在合成与真实世界数据集上的实验证明了基于扩散、无需对齐的视频超分辨率方法的可行性。

原文摘要 · Abstract (English)

In this work, we rethink the approach to video super-resolution by introducing a method based on the Diffusion Posterior Sampling framework, combined with an unconditional video diffusion transformer operating in latent space. The video generation model, a diffusion transformer, functions as a space-time model. We argue that a powerful model, which learns the physics of the real world, can easily handle various kinds of motion patterns as prior knowledge, thus eliminating the need for explicit estimation of optical flows or motion parameters for pixel alignment. Furthermore, a single instance of the proposed video diffusion transformer model can adapt to different sampling conditions without re-training. Empirical results on synthetic and real-world datasets illustrate the feasibility of diffusion-based, alignment-free video super-resolution.

视频超分扩散模型无对齐潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。