arXiv:2509.26325cs.CV2025-09被引 3

用三维傅里叶场统一建模视频时空,实现无伪影超分辨率。

Continuous Space-Time Video Super-Resolution with 3D Fourier Fields

  • 将视频表示为连续时空傅里叶场,避免分步处理与变形补偿
  • 在多种缩放倍数下均实现更清晰、时序更一致的重建效果
  • 适合需要高质量视频增强与高效计算的场景

我们提出一种新的连续时空视频超分辨率框架。不同于将视频分解为独立空间与时间成分并依赖易出错的显式帧对齐进行运动补偿的方法,我们采用连续、时空一致的三维视频傅里叶场(3D Video Fourier Field, VFF)来编码视频。该表示具有三大优势:(1) 可在任意时空位置低成本灵活采样;(2) 同时捕捉精细空间细节与平滑时序动态;(3) 支持解析式高斯点扩散函数,确保任意尺度下的无混叠重建。通过大型时空感受野的神经编码器,从低分辨率输入视频预测傅里叶基的系数。大量实验表明,联合建模显著提升空间与时间超分辨率性能,在多个基准上达到新最佳表现:在广泛缩放因子下,重建结果更锐利、时序更连贯,且计算效率更高。

原文摘要 · Abstract (English)

We introduce a novel formulation for continuous space-time video super-resolution. Instead of decoupling the representation of a video sequence into separate spatial and temporal components and relying on brittle, explicit frame warping for motion compensation, we encode video as a continuous, spatio-temporally coherent 3D Video Fourier Field (VFF). That representation offers three key advantages: (1) it enables cheap, flexible sampling at arbitrary locations in space and time; (2) it is able to simultaneously capture fine spatial detail and smooth temporal dynamics; and (3) it offers the possibility to include an analytical, Gaussian point spread function in the sampling to ensure aliasing-free reconstruction at arbitrary scale. The coefficients of the proposed, Fourier-like sinusoidal basis are predicted with a neural encoder with a large spatio-temporal receptive field, conditioned on the low-resolution input video. Through extensive experiments, we show that our joint modeling substantially improves both spatial and temporal super-resolution and sets a new state of the art for multiple benchmarks: across a wide range of upscaling factors, it delivers sharper and temporally more consistent reconstructions than existing baselines, while being computationally more efficient. Project page: https://v3vsr.github.io.

视频超分傅里叶场时空建模连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。