arXiv:2505.04657eess.IVcs.MM2025-05CVPR被引 14

用事件流提升视频超分的精度、速度和泛化能力

EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events

  • 结合事件流与视频帧,学习长时运动轨迹实现自适应插值
  • 在合成与真实数据集上均超越现有方法,跨尺度泛化更强
  • 适合需要高动态范围和任意分辨率/帧率生成的场景

连续时空视频超分辨率(C-STVSR)旨在同时在任意空间和时间尺度上提升视频质量,近年受到广泛关注。然而,现有方法在分布外的空间与时间尺度上表现不佳。事件流具有高时间分辨率和高动态范围,为视觉任务带来潜力。本文提出EvEnhancer,通过融合事件流的优势,显著提升C-STVSR的有效性、效率与泛化能力。核心包含两部分:1)事件自适应合成利用帧与事件间的时空相关性,捕捉并学习长期运动轨迹,实现信息丰富时空特征的自适应插值与融合;2)局部隐式视频变换器将局部隐式视频神经函数与跨尺度时空注意力结合,学习连续视频表示,以生成任意分辨率与帧率的合理视频。实验表明,EvEnhancer在合成与真实数据集上均优于当前最优方法,且在分布外尺度下具备更优泛化能力。代码已开源:https://github.com/W-Shuoyan/EvEnhancer。

原文摘要 · Abstract (English)

Continuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal scales. On the other hand, event streams characterized by high temporal resolution and high dynamic range, exhibit compelling promise in vision tasks. This paper presents EvEnhancer, an innovative approach that marries the unique advantages of event streams to elevate effectiveness, efficiency, and generalizability for C-STVSR. Our approach hinges on two pivotal components: 1) Event-adapted synthesis capitalizes on the spatiotemporal correlations between frames and events to discern and learn long-term motion trajectories, enabling the adaptive interpolation and fusion of informative spatiotemporal features; 2) Local implicit video transformer integrates local implicit video neural function with cross-scale spatiotemporal attention to learn continuous video representations utilized to generate plausible videos at arbitrary resolutions and frame rates. Experiments show that EvEnhancer achieves superiority on synthetic and real-world datasets and preferable generalizability on out-of-distribution scales against state-of-the-art methods. Code is available at https://github.com/W-Shuoyan/EvEnhancer.

视频超分事件流时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。